<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jcc</journal-id>
      <journal-title-group>
        <journal-title>Journal of Computer and Communications</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-5227</issn>
      <issn pub-type="ppub">2327-5219</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jcc.2026.142009</article-id>
      <article-id pub-id-type="publisher-id">jcc-149820</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>The Financial Digital Divide in the Social Media Era: A Cross-Language Comparative Study of FinBERT Based on Chinese and English Platforms</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Jin</surname>
            <given-names>Xinyi</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> School of Mathematics and Applied Mathematics, Zhejiang Normal University, Jinhua, China </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The author declares no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>10</day>
        <month>02</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>02</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>02</issue>
      <fpage>183</fpage>
      <lpage>212</lpage>
      <history>
        <date date-type="received">
          <day>02</day>
          <month>02</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>25</day>
          <month>02</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>28</day>
          <month>02</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jcc.2026.142009">https://doi.org/10.4236/jcc.2026.142009</self-uri>
      <abstract>
        <p>In order to reveal the manifestations and mechanisms of the cross-language and cross-platform financial digital divide in the social media era, this study is supported by the theory of digital divide, media richness and platform affordances, selects 35,000 financial texts from 10 Chinese and English platforms, uses the unified fine-tuned bilingual FinBERT model combined with statistical testing, and conducts empirical research through progressive hypothesis verification. The research ensures comparability through unified corpus fine-tuning and cross-language alignment, and systematically tests the effects of language, platform, and their interaction. The results show that: Chinese users have lower financial terminology coverage, higher semantic ambiguity, and a gap in expression ability; algorithmic platforms are more likely to disseminate highly emotional and low-professional content, forming an information quality gap; language and platform interact significantly, and language itself is an independent influencing factor of the cognitive empowerment gap. This study expands the cross-linguistic research perspective on the financial digital divide, improves the quantitative measurement method of literacy, provides empirical support for platform optimization, financial science popularization and inclusive finance policy formulation, and points out the research limitations and follow-up directions.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Social Media</kwd>
        <kwd>Financial Digital Divide</kwd>
        <kwd>FinBERT Model</kwd>
        <kwd>Cross-Language Comparison</kwd>
        <kwd>Platform Affordances</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <sec id="sec1dot1">
        <title>1.1. Research Background</title>
        <p>The deep integration of digital finance and social media has reshaped the financial information dissemination ecosystem. The decentralized nature of social media breaks the information monopoly of traditional financial institutions, allows the general public to easily obtain, produce and disseminate financial information, and promotes the popularization of financial knowledge into the era of universal participation [<xref ref-type="bibr" rid="B1">1</xref>]. However, openness has not eliminated information inequality. Instead, it has given rise to a new cross-language and cross-platform financial digital divide due to differences in language systems and heterogeneity of platform mechanisms [<xref ref-type="bibr" rid="B2">2</xref>].</p>
        <p>The “financial digital divide” that this study focuses on specifically refers to the second divide in the digital divide theory that focuses on skills and literacy, and the third divide that focuses on results and empowerment. These two gaps are becoming more and more prominent in the financial field. Although some groups have access to digital financial services, they are unable to effectively interpret and use financial information due to lack of professional literacy, and are unable to achieve cognitive enhancement and welfare improvement, which has become a key bottleneck for the inclusive development of digital finance [<xref ref-type="bibr" rid="B3">3</xref>]. The essential differences in language logic, algorithm design, and functional affordances between Chinese and English platforms have exacerbated the complexity of the gap, and existing research lacks a systematic discussion of the issue of financial digital inequality across contexts [<xref ref-type="bibr" rid="B4">4</xref>].</p>
        <p>With the rapid iteration of financial technology, social media has become a core variable affecting investor decision-making, risk perception, and financial behavior [<xref ref-type="bibr" rid="B5">5</xref>]. However, there are significant differences in the ability of different language groups to use social media financial information. This difference is closely related to the completeness of the financial terminology system, the accuracy of semantic transmission, and the orientation of the platform’s information screening mechanism [<xref ref-type="bibr" rid="B6">6</xref>]. Existing research has limitations such as single scenarios, subjective and extensive measurement methods, and insufficient technical applications [<xref ref-type="bibr" rid="B7">7</xref>].</p>
        <p>To this end, this study [<xref ref-type="bibr" rid="B8">8</xref>] introduces the FinBERT model [<xref ref-type="bibr" rid="B9">9</xref>] dedicated to the financial field, and adopts a text-mining framework inspired by established methodologies in financial linguistics (e.g., Stolper &amp; Walter, 2017), utilizing FinBERT to quantify financial digital literacy through semantic complexity and terminology density into quantifiable text features [<xref ref-type="bibr" rid="B10">10</xref>], accurately depict the second and third financial digital divides, make up for the limitations of traditional research, and provide a new perspective for analyzing cross-language financial information inequality [<xref ref-type="bibr" rid="B11">11</xref>].</p>
      </sec>
      <sec id="sec1dot2">
        <title>1.2. Research Objectives and Hypotheses</title>
        <p>1.2.1. Research Objectives</p>
        <p>The core goal of this study is to reveal the manifestations, internal mechanisms and regulatory effects of the cross-language and cross-platform financial digital divide in the social media era, clarify the independent impact of language type on financial expression ability, clarify the role of platform type in shaping financial information quality, and analyze the impact of the interaction mechanism between language and platform on cognitive empowerment. At the same time, build a quantitative measurement system for financial digital literacy based on text mining, provide a reusable methodological framework for cross-language financial digital divide research, and provide empirical support for platform algorithm optimization, financial science popularization practice, and inclusive financial policy formulation.</p>
        <p>1.2.2. Core Research Questions</p>
        <p>In the social media scenario, how do cross-language contexts and platform characteristics interact to shape the second and third financial digital divides? What are their manifestations and intrinsic mechanisms?</p>
        <p>1.2.3. Research Hypothesis</p>
        <p>Based on the digital divide theory [<xref ref-type="bibr" rid="B12">12</xref>], media richness theory and platform affordance perspective [<xref ref-type="bibr" rid="B13">13</xref>], combined with the logic of cross-language financial communication and existing literature gaps, this study proposes a progressive research hypothesis:</p>
        <p>H1 (Language and financial expression ability gap hypothesis): considering the cross-linguistic variations in financial discourse and cognitive expression (cf. Hofstede, 2011; Pan <italic>et al.</italic>, 2020), which suggest that language-specific structures influence the clarity of professional information [<xref ref-type="bibr" rid="B14">14</xref>], when Chinese social media users discuss financial topics, the density of use of financial terms is lower, the semantic expression is more vague, and the ability to express financial information is significantly weaker than that of English users, which is related to the second financial digital divide in a cross-language context. The core observation indicators draw upon the text mining measurement framework established by Stolper and Walter (2017) [<xref ref-type="bibr" rid="B15">15</xref>]. The semantic ambiguity is measured by the token-level predicted entropy mean output by FinBERT. It is based on the cross-platform text verification of the same financial event and mitigates the selective bias through propensity score matching (PSM) [<xref ref-type="bibr" rid="B16">16</xref>].</p>
        <p>H2 (Gap Hypothesis between Platform Affordance and Information Quality): Based on the interactive logic of media richness theory and platform affordance theory [<xref ref-type="bibr" rid="B17">17</xref>], algorithmic platforms (recommendation streams accounting for &gt;70%) are dominated by visibility affordances and are prone to amplify the spread of emotional information; community-led platforms (attention streams accounting for &gt;60%) are dominated by interactive affordances and are more conducive to the dissemination of professional content [<xref ref-type="bibr" rid="B18">18</xref>]. Therefore, platforms dominated by algorithm recommendations are more likely to disseminate highly emotional and low-professional financial content, forming an information quality gap and exacerbating the second financial digital divide [<xref ref-type="bibr" rid="B19">19</xref>]. The FinBERT model is used to measure the emotional polarization index and content professionalism, and the language type is controlled to ensure that the conclusion is pertinent [<xref ref-type="bibr" rid="B20">20</xref>].</p>
        <p>H3 (cognitive empowerment gap hypothesis under language-platform interaction): After controlling for platform type and selectivity bias, English users are more likely to produce highly empowering content, and language itself is an independent factor in the third financial digital divide. Platform type moderates the strength of the impact of language differences—community-led platforms can mitigate cross-language empowerment gaps through strong interactivity, while algorithmic platforms can amplify this gap [<xref ref-type="bibr" rid="B21">21</xref>]. By constructing four groups of scenarios, interactive analysis was conducted to reveal the synergistic mechanism between the two.</p>
      </sec>
      <sec id="sec1dot3">
        <title>1.3. Research Significance</title>
        <p>1.3.1. Theoretical Significance</p>
        <p>First, expand the application scenarios and measurement dimensions of the financial digital divide theory. Focusing on the second and third divides in the financial field, relying on the text mining measurement protocols of Stolper and Walter (2017) [<xref ref-type="bibr" rid="B22">22</xref>], the text features such as financial term density and semantic ambiguity are transformed into objective quantitative indicators of financial literacy, expanding the empirical boundaries of the cross-language financial digital divide, systematically testing the language-platform interaction effect for the first time, and conducting methodological verification of the text mining measurement method in cross-language scenarios [<xref ref-type="bibr" rid="B23">23</xref>]. Second, deepen the application value of the theoretical integration framework. Systematically integrates three major theories, reveals their interaction in cross-language financial communication scenarios [<xref ref-type="bibr" rid="B24">24</xref>], provides support for the innovative application of traditional theories in the digital age, and responds to the cross-platform research gap proposed by Fjellstrom (2022) [<xref ref-type="bibr" rid="B25">25</xref>]. Third, extend the social science application boundaries of the FinBERT model [<xref ref-type="bibr" rid="B26">26</xref>]. Through unified fine-tuning and cross-language embedding alignment [<xref ref-type="bibr" rid="B27">27</xref>], it is combined with text mining measurement methods and language cognitive theory to improve the methodological system of financial digital inequality research and provide reusable technical paths for similar research [<xref ref-type="bibr" rid="B28">28</xref>]. It should be noted that there may still be limitations in the absolute comparability of cross-language model output, which will be further analyzed later.</p>
        <p>1.3.2. Practical Significance</p>
        <p>For platform operations, it can provide empirical evidence for algorithm optimization. It is recommended that algorithmic platforms balance visibility and interactivity affordances, and refer to Twitter’s Community Notes function to strengthen professional content exposure. For financial science popularization work, it can be targeted to make up for the lack of popularization of Chinese financial terminology and ambiguous semantic transmission, create fragmented and systematic science popularization content, and improve the financial literacy of Chinese users. For cross-border financial information services and policy formulation, the cross-language content conversion mechanism can be optimized to promote the fair flow of global financial information; policymakers can incorporate the cross-language financial digital divide into the inclusive financial policy system, standardize platform content distribution mechanisms, and combat the spread of misleading financial information.</p>
      </sec>
    </sec>
    <sec id="sec2">
      <title>2. Literature Review</title>
      <sec id="sec2dot1">
        <title>2.1. Overview of Core Theoretical Foundations</title>
        <p>2.1.1. Digital Divide Theory and Its Application in the Financial Field</p>
        <p>The digital divide theory has gone through a three-stage evolution of “access-usage-outcomes”, with Van Dijk’s (2006) three-divide framework as the core support: the first focuses on access differences, the second focuses on skill and literacy stratification, and the third reflects cognitive empowerment and welfare gaps. Existing research in the financial field mostly focuses on the first divide. For example, Lu <italic>et al.</italic> (2023) explored the impact of Internet infrastructure on rural financial accessibility. However, empirical tests on the second and third divides are relatively scarce, and research methods have limitations.</p>
        <p>Current relevant research mostly relies on the subjective questionnaire evaluation system of Lusardi and Mitchell (2014), which has the problem of being highly subjective and difficult to capture dynamic expression ability. Only a few studies, such as Aissaoui (2022), have attempted to conduct objective analysis based on user digital behavior data, but they have not broken through the limitations of a single language and a single scenario, making it difficult to accurately capture the dynamic evolution characteristics of the financial digital divide in the algorithmic era.</p>
        <p>There is controversy in the academic community on “whether text features can represent financial literacy”: Amaral &amp; Kolsarici (2020) believe that text expression is only a superficial phenomenon and cannot reflect core capabilities; and Extant research has demonstrated a significant positive correlation between text-based linguistic features and the level of cognitive empowerment in financial contexts. Empirical evidence suggests that users’ ability to articulate complex financial concepts is a reliable proxy for their underlying financial literacy (cf. Hansen <italic>et al.</italic>, 2018) [<xref ref-type="bibr" rid="B29">29</xref>], which can accurately reflect users’ financial expression ability and cognitive level. This study adopts its measurement logic. Semantic ambiguity is measured by the mean token-level predicted entropy output by FinBERT. At the same time, propensity score matching (PSM) is used to alleviate selective bias and improve the objective measurement path of financial digital literacy.</p>
        <p>2.1.2. Integration of Media Richness Theory and Platform Affordances</p>
        <p>Media richness theory was proposed by Daft and Lengel (1986). The core point is that different media have different abilities to convey complex information. High-rich media are more likely to convey information with high ambiguity and complexity, providing a basic framework for analyzing the effect of financial information communication. However, this theory was born in the traditional media era and is difficult to adapt to the technical characteristics of social media. The platform affordance perspective just fills this gap.</p>
        <p>Bucher and Helmond (2018) pointed out that platform functions (<italic>i</italic><italic>.</italic><italic>e.</italic>, affordances) such as algorithm recommendations and interactive mechanisms will directly affect users’ information acquisition behavior and cognitive paths. The core can be summarized into three major dimensions: visibility, interactivity, and connectivity. Although existing research has begun to integrate two major theories to analyze digital information dissemination, such as Burke and Hung (2021) to explore users’ learning and participation behaviors on social platforms, there are still three shortcomings: First, the research scenario is limited to a single language context, and the essential differences in affordance design of Chinese and English platforms are not compared; second, there is a lack of in-depth analysis of the interaction mechanism; third, there is a lack of pertinence, and the factors regulating the communication effect are not combined with the professional and risk characteristics of financial information. This gap was also mentioned in the cross-platform research of Kim <italic>et al.</italic> (2021).</p>
        <p>It should be clear that platform affordances are not inherent properties of the platform, but emergent features formed by the interaction between users and technology (Bucher &amp; Helmond, 2018). Therefore, the dichotomy of “algorithm-led/community-led” in this study is a heuristic operationalization based on core functional features, aiming to simplify complex interactive relationships and focus on the impact of platform-led mechanisms on financial information quality. There is a complex interactive tension between media richness and platform affordances: high affordances do not equal high richness. Algorithmic platforms have strong visibility and affordances (recommendation streams account for &gt;70%), but the content forms are mostly short videos and short texts, with low media richness and difficulty in carrying complex financial logic; community-led platforms are dominated by interactive affordances (attention streams account for &gt;60%), and content forms such as long texts and in-depth discussions have higher media richness, which is more conducive to the transmission of professional financial information. This interaction logic is the core theoretical basis of H2.</p>
        <p>2.1.3. Theoretical Basis of Cross-Language Financial Communication</p>
        <p>Cross-cultural linguistics and financial communication research shows that there are structural differences in the Chinese and English financial semantic expression systems (cf. Pan <italic>et al.</italic>, 2020) [<xref ref-type="bibr" rid="B30">30</xref>], which are shaped by multiple factors such as the development history of the financial market, language expression habits, and the perfection of the science popularization system, which directly affect the cognitive empowerment effect. From the perspective of language cognitive theory, this difference is reflected in three aspects: First, the difference in terminology system. English financial terminology system is precise and has high penetration rate, while Chinese terminology is mostly derived from translation, has semantic overlap and ambiguity problems, and has low popularity; second, is the difference in expression habits. English financial communication focuses on logic and data support, while Chinese is biased towards experience sharing and emotional expression. Third, there is a difference in science popularization systems. English financial science popularization starts early and has wide coverage, while Chinese financial science popularization mainly focuses on fragmented content. This structural difference is a core driver of the cross-language financial digital divide.</p>
        <p>At the same time, Chu <italic>et al.</italic> (2022) pointed out that there are high-quality professional content communities such as Snowball in Chinese social media, but most of their users are subgroups with high financial literacy, which suffers from selective bias and are difficult to represent the overall Chinese users. This study controls the differences between professional users and ordinary user groups through user stratification screening and analysis in data cleaning, ensuring that the research conclusions reflect the overall characteristics of different language groups, and also provides a reference for testing the interaction between language and platform.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Research on Social Media Financial Information Dissemination and Digital Inequality</title>
        <p>The decentralized nature of social media has reconstructed the financial information dissemination model. Gupta and Chen (2020) confirmed that it broke the information monopoly of traditional financial institutions and enabled ordinary users to become core producers and disseminators. However, it did not achieve information equality and instead gave birth to a new type of digital inequality. Existing research has four limitations.</p>
        <p>First, the research scenario is single and lacks a cross-language comparative perspective. Although Chu <italic>et al.</italic> (2022) found the impact of cross-language differences, it was limited by a small sample; second, the measurement method is subjective and extensive, and the evaluation framework of Huang <italic>et al.</italic> (2023) lacks precise technical support. This study supplements the token-levelThe mean value of predicted entropy measures semantic ambiguity; third, the research design has endogeneity risks, the interaction effect is not tested and collinearity issues are not properly handled; fourth, competing explanations are ignored, and the conclusions are not comprehensive and objective.</p>
        <p>The tension in these research conclusions provides a core entry point for this study. Related counterexamples show that platform type may moderate the impact of language differences, confirming the necessity of conducting cross-language and cross-platform interaction research.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Application of FinBERT Model in Financial Text Analysis</title>
        <p>FinBERT is a special pre-training model in the financial field proposed by Araci (2019). It is based on general BERT and fine-tuned and optimized with massive financial texts. It has significant advantages in tasks such as term recognition and semantic understanding. After being optimized by ProsusAI (2019), it is widely used in financial sentiment analysis, risk identification and other scenarios. Liu <italic>et al.</italic> (2020) confirmed that its performance is significantly better than traditional methods and general BERT, with a related improvement of over 20%. Existing applications mostly focus on the financial market level, but there are still three major gaps: application scenarios are limited to technical fields, lack of cross-language comparative research, and weak theoretical integration. This study aims to improve comparability through unified fine-tuning and cross-language embedding alignment, and deeply integrates it with the financial digital divide theory to achieve collaborative advancement of technology and theoretical research.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Literature Review and Research Gaps</title>
        <p>Existing research provides theoretical foundations, methodological references and core entry points, but there are still four core gaps:</p>
        <p>First, the logic of the relationship between indicators and constructs is fuzzy and the measurement basis is insufficiently supported. Existing research has not clarified the theoretical connection between text characteristics and financial literacy. Existing text mining measurement methods (e. g., Stolper &amp; Walter, 2017) have not been fully verified in cross-language scenarios, and the quantitative method of semantic ambiguity has not been clarified. This study supplements relevant measurement methods and improves the cross-language comparability of indicators; the classification of platform types lacks a unified theoretical anchor, which affects the objectivity and comparability of conclusions.</p>
        <p>Second, the depth of theoretical integration is insufficient and there is a lack of research on interaction mechanisms. Existing research has not deeply explored the interaction mechanism of “media richness + platform affordance + cross-language differences”. It is difficult to reveal the complex formation mechanism of the financial digital divide by analyzing the role of a single variable in isolation.</p>
        <p>Third, the research design has endogeneity risks and the credibility of causal inferences is weak. Existing research does not properly handle the problem of collinearity between language and platform, lacks effective control over user heterogeneity and selective bias, makes it difficult to distinguish independent effects, and concludes with insufficient causal explanation.</p>
        <p>Fourth, competing explanations are ignored and the theoretical contribution is insufficiently highlighted. Existing studies mostly support a single conclusion and do not fully respond to counterexamples and competing explanations, showing that the conclusions are not comprehensive enough to form a breakthrough theoretical contribution.</p>
        <p>This study addresses the above gaps by clarifying the logic of correlation between indicators and constructs, deepening the integration of multiple theories, optimizing research design, responding to competing explanations, and building a cross-language and cross-platform empirical analysis framework to make up for the shortcomings of existing research. It also expands the application scenarios and theoretical practice boundaries of the FinBERT model to provide new perspectives and methodological support for financial digital divide research.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Research Methods</title>
      <sec id="sec3dot1">
        <title>3.1. Research Methodological Framework (Based on the Research Onion Model)</title>
        <p>This study follows Saunders <italic>et al.</italic>’s (2019) “Research Onion Model” to construct a complete methodological system. Empiricism is adopted at the philosophical level, focusing on the observable text characteristics of the financial digital divide, and verifying hypotheses with quantitative data; the research paradigm is quantitative research, based on 35,000 pieces of multi-platform text data, using the FinBERT model to extract quantitative features and combining statistical testing to avoid subjective bias; the research strategy combines cross-platform comparison and cross-sectional research to accurately capture the characteristics of the divide; the research method integrates text mining and statistical analysis to ensure the scientificity of the conclusions; data is collected through compliance interfaces, combining FinBERT fine-tuning and PythonTools that take into account both legality and analytical efficiency.</p>
      </sec>
      <sec id="sec3dot2">
        <title>3.2. Data Sources and Processing</title>
        <p>3.2.1. Data Collection Design</p>
        <p>The data for this study was obtained through a collection method based on public APIs and compliant data interfaces, covering 10 Chinese and English social media platforms, taking into account the original platforms and new professional financial platforms, and finally obtained 35,000 pieces of valid data, which was completely consistent with the preset goal.</p>
        <p>The core design of data collection is as follows: First, the platform coverage is comprehensive, covering two dimensions: Chinese/English and algorithmic/community type. The Chinese platform includes algorithmic type (Weibo, Xiaohongshu, Douyin) and community type (Snowball, Oriental Fortune, Flush), and the English platform includes algorithmic type (Twitter, TikTok, Facebook, LinkedIn) and community-based (Reddit) to make up for the limitations of a single platform for existing research; the second is the precise adaptation of the topic library, which expands the topic library based on core topics in the financial field. Chinese topics include 34 keywords (such as “fund”, “dragon and tiger list”, “MACD”), and English topics include 34 corresponding keywords (such as “fund”, “MACD”, “Bollinger Bands”) to ensure that the text focuses on the financial field; third, the data volume is distributed in a balanced manner, and the target data volume is reasonably divided according to platform type and language dimensions to avoid sample bias caused by an excessive proportion of data from a single platform. The specific data distribution is shown in <bold>Table 1</bold> below.</p>
        <p>3.2.2. Data Preprocessing Steps</p>
        <p>In order to ensure data quality, this study is based on Python’s Pandas, Jieba, NLTK and other libraries, and uses the GPU acceleration function on the Colab platform to carry out systematic data cleaning. The whole process takes about 2.5 hours. There are 38,742 pieces of original data, and 35,000 pieces of valid data are</p>
        <p>Table 1. Data distribution on different platforms.</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>language type</td>
                <td>platform type</td>
                <td>Platform name</td>
                <td>Data volume (items)</td>
                <td>Proportion</td>
                <td>Platform feature adaptation instructions</td>
              </tr>
              <tr>
                <td>Chinese (zh)</td>
                <td>Algorithmic</td>
                <td>Weibo</td>
                <td>2500</td>
                <td>7.14%</td>
                <td>Popular financial topic communication, concise text and highly interactive</td>
              </tr>
              <tr>
                <td>Chinese (zh)</td>
                <td>Algorithmic</td>
                <td>little red book</td>
                <td>2500</td>
                <td>7.14%</td>
                <td>Sharing of life-oriented finance, including experience summaries and pitfall avoidance guides</td>
              </tr>
              <tr>
                <td>Chinese (zh)</td>
                <td>Algorithmic</td>
                <td>Tik Tok</td>
                <td>2500</td>
                <td>7.14%</td>
                <td>Short video comment area text, emotional expression is obvious</td>
              </tr>
              <tr>
                <td>Chinese (zh)</td>
                <td>community type</td>
                <td>snowball</td>
                <td>2500</td>
                <td>7.14%</td>
                <td>Professional investor community, focusing on market analysis and strategy sharing</td>
              </tr>
              <tr>
                <td>Chinese (zh)</td>
                <td>community type</td>
                <td>Oriental Fortune</td>
                <td>4000</td>
                <td>11.43%</td>
                <td>Professional financial community, including in-depth content such as Dragon and Tiger List, main funds, etc.</td>
              </tr>
              <tr>
                <td>Chinese (zh)</td>
                <td>community type</td>
                <td>Flush</td>
                <td>4000</td>
                <td>11.43%</td>
                <td>A supporting community for stock trading tools, focusing on indicator analysis and stock selection formulas</td>
              </tr>
              <tr>
                <td>English (en)</td>
                <td>Algorithmic</td>
                <td>Twitter (X)</td>
                <td>3750</td>
                <td>10.71%</td>
                <td>Dissemination of global financial topics with strong real-time nature and diverse viewpoints</td>
              </tr>
              <tr>
                <td>English (en)</td>
                <td>Algorithmic</td>
                <td>TikTok</td>
                <td>3750</td>
                <td>10.71%</td>
                <td>Short video comment area, youthful expression, extreme emotions prominent</td>
              </tr>
              <tr>
                <td>English (en)</td>
                <td>Algorithmic</td>
                <td>Facebook</td>
                <td>3500</td>
                <td>10.00%</td>
                <td>Life-oriented financial discussions, including sharing of personal investment experiences</td>
              </tr>
              <tr>
                <td>English (en)</td>
                <td>Algorithmic</td>
                <td>LinkedIn</td>
                <td>3500</td>
                <td>10.00%</td>
                <td>Workplace finance discussion, focusing on professional frameworks and institutional perspectives</td>
              </tr>
              <tr>
                <td>English (en)</td>
                <td>community type</td>
                <td>Reddit</td>
                <td>2500</td>
                <td>7.14%</td>
                <td>Professional investment community with a lot of in-depth analysis and backtest verification content</td>
              </tr>
              <tr>
                <td>total</td>
                <td>total</td>
                <td>10 platforms</td>
                <td>35,000</td>
                <td>100%</td>
                <td>-</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The data collection fields strictly correspond to the research variable requirements. The core fields include 15 items: text unique ID (text_id), platform identification (platform), language type (language), platform type (platform_type), crawling time (crawl_time), publishing time (post_time), text content (text_content), text length(text_length), hashtags (hashtags), number of likes (like_count), number of comments (comment_count), number of shares (share_count), data quality level (data_quality), clean notes (clean_note) and FinBERT derived indicators (to be added later) to ensure field integrity and analyzability.</p>
        <p>retained after cleaning, with a cleaning rate of 9.65%, which meets the data quality standards for large-sample quantitative research. The cleaning process is as follows: First, double deduplication of “text unique ID + text content” is performed, eliminating 1289 pieces of duplicate data, with a deduplication accuracy of 99.2%; then, by referring to the Chinese and English financial keyword dictionary (286 Chinese, 242 English) constructed by the authoritative manual, financial-related texts are screened, and 1458 irrelevant data such as life and entertainment are eliminated; then, the Chinese and English text characteristics are optimized respectively, and the Chinese text is processed by JiebaWord segmentation removes stop words, special symbols and garbled characters, and the English word form is restored and unified through the NLTK library. The batch processing rate is 1200 items/minute, and the word segmentation accuracy is 93.5%. Then, 69 popular content and 826 short texts are eliminated, and FinBERT indicator outliers are eliminated through Z test (|Z| &gt; 3). Finally, Subword Token is performed based on FinBERT’s own AutoTokenizer. Split, standardize the Chinese and English text according to “50 tokens/information unit” (intercept the first 50 tokens, and complete the remaining ones), and calculate the average coverage rate of financial terms in each unit. This method adapts to the differences in Chinese and English language structures, is closer to the actual semantic processing logic of the model, and improves the comparability of cross-language indicators.</p>
      </sec>
      <sec id="sec3dot3">
        <title>3.3. Operationalization of Core Variables and Model Optimization</title>
        <p>3.3.1. Operational Definition of Variables</p>
        <p>Based on three major research hypotheses, abstract concepts are transformed into quantifiable text features and classification variables, and the FinBERT model and statistical tools are combined to achieve variable measurement. The specific operationalization plan is shown in <bold>Table 2</bold> below to ensure that the variables are clearly defined and the measurement methods are reproducible.</p>
        <p>3.3.2. FinBERT Model Fine-Tuning and Verification</p>
        <p>This research is based on the bilingual FinBERT model optimized by ProsusAI, and is specially fine-tuned for financial text analysis and cross-language comparison needs to improve the accuracy of term recognition, semantic understanding and emotional polarization judgment. The fine-tuning process is: construct a Chinese and English annotation data set (3000 Chinese items, 2000 English items, Cohen’s Kappa = 0.82) (The annotation team consists of 3 bilingual financial experts (all with over 5 years of financial work experience and TEM-8 certification). The annotation process follows three steps: ① Jointly annotate 100 texts to unify annotation standards (refer to the empowerment dimension definition in Appendix A); ② Independently annotate the remaining 4900 texts (3000 Chinese, 2000 English); ③ Discrepancy handling: For texts with inconsistent annotations (Kappa &lt; 0.7), final annotations were reached through tripartite consultation, ensuring the overall Cohen’s Kappa = 0.82 (high consistency).</p>
        <p>Table 2. Variable operationalization scheme.</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>correspondence hypothesis</td>
                <td>Variable type</td>
                <td>variable name</td>
                <td>operational definition</td>
                <td>Measurement methods and tools</td>
                <td>Variable value range</td>
              </tr>
              <tr>
                <td>H1</td>
                <td>independent variable</td>
                <td>language type</td>
                <td>Categorical variables to distinguish Chinese from English text</td>
                <td>Extract the collection field “language”, Chinese is assigned a value of 1, and English is assigned a value of 0</td>
                <td>0 (English), 1 (Chinese)</td>
              </tr>
              <tr>
                <td>H1</td>
                <td>dependent variable</td>
                <td>financial terminology coverage</td>
                <td>Proportion of financial terms to the total number of valid words in standardized texts</td>
                <td>FinBERT extracts terms + keyword dictionary calibration, and calculates the proportion after normalizing the text by “50 tokens/information unit”</td>
                <td>0 - 1 (the larger the number, the higher the term density)</td>
              </tr>
              <tr>
                <td>H1</td>
                <td>dependent variable</td>
                <td>Semantic uncertainty score</td>
                <td>The degree of semantic ambiguity of core financial vocabulary (the higher the entropy value, the more ambiguous it is)</td>
                <td>1. Build a bilingual financial core dictionary and double-verify to extract core words; 2. Process the core words [MASK] to calculate the mean entropy value; 3. Min-max normalization mapping to the 0 - 10 interval</td>
                <td>0 - 10 (the larger the value, the more ambiguous the semantics)</td>
              </tr>
              <tr>
                <td>H2</td>
                <td>independent variable</td>
                <td>platform type</td>
                <td>Categorical variables distinguishing algorithm-led from community-led platforms</td>
                <td>Based on Bucher &amp; Helmond’s (2018) platform affordance theory (affordances are emergent attributes of the user-technology relationship, not inherent characteristics of the platform), combined with SimilarWeb traffic data, “recommended flow proportion &gt; 70%” is defined as algorithm- dominated (1), and “following flow proportion &gt; 60%” is defined as community-dominated (0). This dichotomy is a heuristic operationalization to capture the dominant characteristics of the platform, rather than an absolute classification.</td>
                <td>0 (community type), 1 (algorithm type)</td>
              </tr>
              <tr>
                <td>H2</td>
                <td>dependent variable</td>
                <td>emotional polarization index</td>
                <td>The extent to which the text expresses extreme positive and negative emotions(capturing two-tailed polarization to avoid single-dimensional bias)</td>
                <td>① FinBERT outputs sentiment scores (range: -1~1); ② Set two-tailed thresholds: texts with sentiment scores &gt; 0.7 or &lt; −0.7 are defined as “polarized texts” (thresholds calibrated based on the 90th percentiles of positive/negative sentiment scores in 5000 annotated texts); ③ Calculate the proportion of polarized texts as the index value. See Appendix B for detailed calibration process.</td>
                <td>0 - 1 (the larger the value, the higher the degree of polarization)</td>
              </tr>
              <tr>
                <td>H2</td>
                <td>dependent variable</td>
                <td>Content professionalism score</td>
                <td>Financial professionalism and informational value of the text</td>
                <td>FinBERT classifier trained on 5000 annotated data, scored from conceptual accuracy, data completeness, and logical interpretability</td>
                <td>0 - 1 (the larger the value, the higher the professionalism)</td>
              </tr>
              <tr>
                <td>H3</td>
                <td>independent variable</td>
                <td>Language- Platform interactions</td>
                <td>Intersection variables of language and platform type (reflecting joint effects)</td>
                <td>Language type (0/1) × platform type (0/1), generating four types of interaction combinations</td>
                <td>00, 01, 10, 11 (corresponding to four types of scenarios)</td>
              </tr>
              <tr>
                <td>H3</td>
                <td>moderator variable</td>
                <td>platform type</td>
                <td>Same as H2 independent variable (used to control individual effects)</td>
                <td>Same as H2 measurement method, algorithm type is assigned a value of 1, community type is assigned a value of 0</td>
                <td>0 (community type), 1 (algorithm type)</td>
              </tr>
              <tr>
                <td>H3</td>
                <td>dependent variable</td>
                <td>Financial Cognition Empowerment Score</td>
                <td>The extent to which the text contains enabling content such as explanation mechanisms and risk warnings</td>
                <td>FinBERT fine-tunes model classification, calculates the proportion of enabling content and semantic contribution (including the three dimensions of mechanism, risk, and decision-making)</td>
                <td>0 - 1 (the larger the value, the stronger the empowerment effect)</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Detailed calculation process of the semantic uncertainty score: ① Selection of core financial vocabulary: Based on A Dictionary of English-Chinese Financial Terms and Encyclopedia of China Finance, the top 30% most frequently occurring core vocabulary are selected (200 terms each for Chinese and English, one-to-one corresponding, e.g., “MACD—指数平滑异同移动平均线”); ② Masking strategy: Each core vocabulary in the text is replaced with [MASK] one by one without repeated masking in a single text; ③ Entropy calculation: Only for the masked core vocabulary, extract the token-level predicted entropy output by FinBERT, and take the average of the entropy values of all core vocabulary; ④ Normalization: Map to the 0 - 10 interval through min-max normalization (original entropy value range: 0 - 3.2).</p>
        <p>The training and evaluation protocol for the professionalism and empowerment classifiers is specified as follows: ① Dataset split: The 5000 annotated texts are divided into a training set (3500 texts), a validation set (500 texts), and a test set (1000 texts) at a 7:1:2 ratio; ② Hyperparameters: Learning rate = 2e−5, batch size = 32, training epochs = 5, Dropout probability = 0.1, with an early stopping strategy (patience = 2, based on the validation set F1 score) to avoid overfitting; ③ Class balance: The SMOTE technique is used to oversample the minority class (highly empowering content), making the class ratio in the training set close to 1:1; ④ Evaluation metrics: On the test set, the professionalism classifier achieves Accuracy = 0.87, Precision = 0.85, Recall = 0.83, and F1 = 0.84; the empowerment classifier achieves Accuracy = 0.86, Precision = 0.82, Recall = 0.80, and F1 = 0.81, indicating stable model performance.</p>
        <p>Set core parameters based on the Colab GPU environment and introduce DropoutThe layer avoids overfitting and is adapted to cross-language scenarios through bilingual alignment training; the verification results show that the term recognition accuracy is 92.3%, the F1 values of semantic uncertainty judgment and emotional polarization classification are 0.88 and 0.86 respectively, and the empowerment score prediction error is &lt;5%. “Token-level prediction entropy” is an indicator of semantic ambiguity. Its prediction uncertainty stems from insufficient semantic support of the text. After training with the same annotated data and manual verification of 200 texts (correlation coefficient 0.79), it can effectively reflect differences in users’ expressive abilities.</p>
        <p>This study adopts a dual-branch fusion architecture: the base model takes ProsusAI FinBERT-base (English pre-trained) as the core, integrated with Chinese BERT-base (developed by the Harbin Institute of Technology and iFLYTEK Joint Laboratory) as the Chinese processing branch to ensure adaptability for Chinese semantic understanding. The cross-linguistic strategy employs a “MUSE tool + bilingual financial corpus alignment” scheme: first, a bilingual financial parallel corpus is constructed (including 50,000 Chinese-English financial dictionary entries, 10,000 Chinese-English financial report summaries, and 3000 Chinese-English financial policy translations); then, supervised alignment training via the MUSE tool maps the embedding spaces of Chinese and English models into a unified 768-dimensional vector space. The cross-language embedding alignment steps are: 1) Pre-train a cross-linguistic mapping matrix using the bilingual corpus; 2) Perform linear transformation on the token embeddings of Chinese and English outputs from FinBERT; 3) Unify vector norms through L2 normalization to ultimately achieve comparable cross-language embeddings. After model fine-tuning, the cross-linguistic consistency accuracy of Chinese and English terminology recognition reaches 89.7%.</p>
        <p>Validity verification of the semantic uncertainty metric: 500 Chinese and English texts (250 each) were selected, and independently scored by 2 bilingual financial experts using a “1 - 5 point ambiguity scale” (1 = completely clear semantics, 5 = extremely ambiguous semantics). Discrepancies were resolved through third-party expert arbitration (Cohen’s Kappa = 0.81). The Pearson correlation coefficient between the metric entropy value and expert scores was r = 0.79 (p &lt; 0.001), confirming that the metric effectively reflects human-judged semantic ambiguity.</p>
      </sec>
    </sec>
    <sec id="sec4">
      <title>4. Research Results and Data Analysis</title>
      <sec id="sec4dot1">
        <title>4.1. Descriptive Statistical Analysis</title>
        <p>Based on the research scope and data collection plan defined above, this study obtained 35,000 pieces of valid social media financial text data, and formed standardized analysis samples after missing value removal, outlier filtering and text cleaning. Focusing on core variables such as financial term coverage, semantic uncertainty score, and emotional polarization index, the mean, standard deviation, extreme value and other indicators are calculated through Python’s Scikit-learn library, and visual charts such as kernel density charts and word cloud charts are generated using the Colab platform to initially present variable distribution characteristics and group differences, laying the foundation for subsequent hypothesis testing. Combined with the three major research hypotheses, descriptive statistics focus on the grouping differences of language type and platform type. The core variable statistical results are shown in <bold>Table 3</bold> below. This study uses 50 tokens to standardize Chinese and English texts. The average number of tokens for Chinese samples is 48.2 and English 51.7. The distribution is balanced, ensuring the equivalence of cross-language comparisons.</p>
        <p>In order to ensure cross-language comparability, this study uses FinBERT’s Subword Token mechanism to standardize text (instead of counting character counts) and calculate core indicators based on 50 tokens/information unit. According to statistics, the average number of tokens in Chinese samples is 48.2 (standard deviation 6.3), and the average number of tokens in English samples is 51.7 (standard deviation 5.8). The distribution of the two is balanced and there is no significant difference (t = 1.82, p &gt; 0.05), ensuring the equivalence of comparison benchmarks for indicators such as term coverage. The description of differences grouped by language and platform type is in the same direction as the previous theoretical hypothesis: in terms of language dimension, the average coverage rate of financial terms in English texts (0.24) is significantly higher than that in Chinese (0.12), and the average semantic uncertainty score (3.85) is significantly lower than that in Chinese (5.59)., preliminarily confirming the gap in cross-language financial expression capabilities; in terms of platform dimension, the content professionalism score of community-led platforms (4.68) is higher than that of algorithmic platforms (2.03), and the emotional polarization index (0.32) is lower than algorithmic platforms (0.49), which is in line with the expectations of platform affordance theory.</p>
        <p>Table 3. Statistical results of core variables.</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>variable name</td>
                <td>Sample size (strips)</td>
                <td>mean</td>
                <td>standard deviation</td>
                <td>minimum value</td>
                <td>maximum value</td>
                <td>Description of data characteristics</td>
              </tr>
              <tr>
                <td>financial terminology coverage</td>
                <td>35,000</td>
                <td>0.18</td>
                <td>0.09</td>
                <td>0.02</td>
                <td>0.63</td>
                <td>The overall coverage rate is low and individual differences are significant, which is in line with the fragmented expression characteristics of social media.</td>
              </tr>
              <tr>
                <td>Semanticuncertainty score</td>
                <td>35,000</td>
                <td>4.72</td>
                <td>1.85</td>
                <td>0.83</td>
                <td>8.96</td>
                <td>Moderate to above, semantic ambiguity is common, consistent with the complexity of financial information</td>
              </tr>
              <tr>
                <td>emotional polarization index</td>
                <td>35,000</td>
                <td>0.41</td>
                <td>0.22</td>
                <td>0.05</td>
                <td>0.92</td>
                <td>Some texts have prominent extreme emotions, confirming the emotional transmission characteristics of social media</td>
              </tr>
              <tr>
                <td>Content professionalism score</td>
                <td>35,000</td>
                <td>3.26</td>
                <td>1.43</td>
                <td>1.00</td>
                <td>8.50</td>
                <td>1 - 10 point scale, overall professionalism is weak, in line with the positioning of mass communication scenarios</td>
              </tr>
              <tr>
                <td>Proportion of highly empowering content</td>
                <td>35,000</td>
                <td>0.23</td>
                <td>0.42</td>
                <td>0.00</td>
                <td>1.00</td>
                <td>Dichotomous variable (0 = no, 1 = yes), the overall proportion of high-enabling content is low</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Visual analysis further strengthens the above characteristics: the kernel density plot (<xref ref-type="fig" rid="fig1">Figure 1</xref>) shows that the distribution curves of financial terminology coverage in Chinese and English texts are clearly separated, and the English text peaks tend to be in high segments and have lower dispersion; the word cloud chart (<xref ref-type="fig" rid="fig2">Figure 2</xref>) shows that English texts are mostly standardized terms, while Chinese texts are mainly expressed in daily life, and professional terms are limited to basic vocabulary, which provides visual support for the research hypothesis.</p>
        <fig id="fig1">
          <label>Figure 1</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId13.jpeg?20260228085947" />
        </fig>
        <p>Figure 1. Kernel density distribution chart of Chinese and English financial terminology coverage.</p>
        <fig id="fig2">
          <label>Figure 2</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId14.jpeg?20260228085947" />
        </fig>
        <p>Figure 2. Comparison chart of Chinese and English financial text word clouds.</p>
      </sec>
      <sec id="sec4dot2">
        <title>4.2. Hypothesis Test Results</title>
        <p>4.2.1. H1 Test: Gap between Language and Financial Expression Skills</p>
        <p>In order to verify H1 (there is a significant gap in financial expression ability between Chinese and English), this study uses language type as the independent variable, financial term coverage and semantic uncertainty score as the dependent variables, and uses the independent sample t test combined with the propensity score matching (PSM) method (taking text length as matching variables) for analysis. The results show that there are significant differences between Chinese and English in the two major dependent variables (p &lt; 0.001): the average financial term coverage rate of English users (0.24 ± 0.08) is higher than that of Chinese (0.12 ± 0.07), and the semantic uncertainty score of Chinese users (5.59 ± 1.72) is higher than that of English (3.85 ± 1.53). The effect sizes are all large. After PSM matching (sample 28,000 items), the differences are still significant and the results are stable. This shows that the structural differences in the Chinese and English financial semantic expression systems and the institutional environment (such as the Chinese financial education system and science popularization system) work together to create a significant financial expression ability gap among Chinese users, which is related to the second financial digital divide in a cross-language context.</p>
        <p>Propensity Score Matching (PSM) adopts kernel matching, with matching variables including text length plus 4 additional covariates: ① Platform type (0 = community-led, 1 = algorithm-led); ② Topic keyword clusters (texts classified into 5 categories via LDA clustering: funds, stocks, bonds, macroeconomics, derivatives, coded as 0 - 4); ③ Posting time window (±7 days to ensure consistent time effects); ④ Engagement proxy variable (logarithmic transformation of likes + comments to control user activity differences). The final sample size after matching is 28,000 texts (14,000 each for Chinese and English) (<bold>Table 4</bold>).</p>
        <p>Table 4. PSM matching balance diagnostic table.</p>
        <table-wrap id="tbl4">
          <label>Table 4</label>
          <table>
            <tbody>
              <tr>
                <td>Covariates</td>
                <td>SMD before Matching</td>
                <td>SMD after Matching</td>
                <td>Balance Improvement</td>
              </tr>
              <tr>
                <td>Text Length</td>
                <td>0.08</td>
                <td>0.03</td>
                <td>Meets standard (SMD &lt; 0.1)</td>
              </tr>
              <tr>
                <td>Platform Type</td>
                <td>0.21</td>
                <td>0.07</td>
                <td>Meets standard</td>
              </tr>
              <tr>
                <td>Topic Keyword Clusters</td>
                <td>0.19</td>
                <td>0.05</td>
                <td>Meets standard</td>
              </tr>
              <tr>
                <td>Engagement Proxy Variable</td>
                <td>0.25</td>
                <td>0.09</td>
                <td>Meets standard</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Note: A standardized mean difference (SMD) &lt; 0.1 indicates good inter-group balance. All covariates meet this standard after matching, effectively controlling for selection bias. The visual distribution of SMD values of all covariates before and after PSM matching is shown in Appendix C, which more intuitively reflects the significant improvement of in-ter-group balance after matching.</p>
        <p>4.2.2. H2 Test: Platform Affordance and Information Quality Gap</p>
        <p>In order to verify H2 (platform type affects the quality of financial information and creates an information quality gap), this study uses platform type as the independent variable, emotional polarization index and content professionalism score as the dependent variables, and adopts a one-factor analysis of variance (ANOVA) test. At the same time, the language type is controlled as a covariate to eliminate interference. The results show that the main effect of platform type on the two major dependent variables is significant (p &lt; 0.001): the emotional polarization index of algorithmic platforms (0.49 ± 0.21) is higher than that of community-based platforms (0.32 ± 0.18), and the content professionalism score of community-led platforms (4.68 ± 1.35) is significantly higher than that of algorithmic platforms (2.03 ± 1.12). The effect sizes are medium and large effects respectively. This difference is significant in both Chinese and English scenarios (p &lt; 0.01), and is consistent across languages. The boxplot (<xref ref-type="fig" rid="fig3">Figure 3</xref>) intuitively presents: the interquartile range of professionalism scores of community-led platforms is concentrated in high segments and the data is stable, while the emotional polarization index of algorithmic platforms has many outliers and a high median. The results support H2, that is, algorithmic platforms are more likely to form a high-emotional, low-professional information quality gap, exacerbating the second financial digital divide.</p>
        <fig id="fig3">
          <label>Figure 3</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId15.jpeg?20260228085950" />
        </fig>
        <p>Figure 3. Box plot of content professionalism scores for different platform types.</p>
        <p>4.2.3. H3 Test: Cognitive Empowerment Gap under Language-Platform Interaction</p>
        <p>A logistic regression model was used instead of ANOVA for group-level proportion analysis, with “whether a single text is highly empowering content” (0 = no, 1 = yes) as the dependent variable. The model is constructed as follows:</p>
        <disp-formula id="FD1">
          <mml:math display="inline">
            <mml:mtable>
              <mml:mtr>
                <mml:mtd>
                  <mml:mi>log</mml:mi>
                  <mml:mi>i</mml:mi>
                  <mml:mi>t</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>P</mml:mi>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:mi>H</mml:mi>
                          <mml:mi>i</mml:mi>
                          <mml:mi>g</mml:mi>
                          <mml:mi>h</mml:mi>
                          <mml:mi>l</mml:mi>
                          <mml:mi>y</mml:mi>
                          <mml:mi>E</mml:mi>
                          <mml:mi>m</mml:mi>
                          <mml:mi>p</mml:mi>
                          <mml:mi>o</mml:mi>
                          <mml:mi>w</mml:mi>
                          <mml:mi>e</mml:mi>
                          <mml:mi>r</mml:mi>
                          <mml:mi>i</mml:mi>
                          <mml:mi>n</mml:mi>
                          <mml:mi>g</mml:mi>
                          <mml:mo>=</mml:mo>
                          <mml:mn>1</mml:mn>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>=</mml:mo>
                  <mml:mi>α</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:msub>
                    <mml:mi>β</mml:mi>
                    <mml:mn>1</mml:mn>
                  </mml:msub>
                  <mml:mo>×</mml:mo>
                  <mml:mi>L</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>n</mml:mi>
                  <mml:mi>g</mml:mi>
                  <mml:mi>u</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>g</mml:mi>
                  <mml:mi>e</mml:mi>
                  <mml:mi>T</mml:mi>
                  <mml:mi>y</mml:mi>
                  <mml:mi>p</mml:mi>
                  <mml:mi>e</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:msub>
                    <mml:mi>β</mml:mi>
                    <mml:mn>2</mml:mn>
                  </mml:msub>
                  <mml:mo>×</mml:mo>
                  <mml:mi>P</mml:mi>
                  <mml:mi>l</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>t</mml:mi>
                  <mml:mi>f</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>m</mml:mi>
                  <mml:mi>T</mml:mi>
                  <mml:mi>y</mml:mi>
                  <mml:mi>p</mml:mi>
                  <mml:mi>e</mml:mi>
                </mml:mtd>
              </mml:mtr>
              <mml:mtr>
                <mml:mtd>
                  <mml:mo>+</mml:mo>
                  <mml:msub>
                    <mml:mi>β</mml:mi>
                    <mml:mn>3</mml:mn>
                  </mml:msub>
                  <mml:mo>×</mml:mo>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mi>L</mml:mi>
                      <mml:mi>a</mml:mi>
                      <mml:mi>n</mml:mi>
                      <mml:mi>g</mml:mi>
                      <mml:mi>u</mml:mi>
                      <mml:mi>a</mml:mi>
                      <mml:mi>g</mml:mi>
                      <mml:mi>e</mml:mi>
                      <mml:mi>T</mml:mi>
                      <mml:mi>y</mml:mi>
                      <mml:mi>p</mml:mi>
                      <mml:mi>e</mml:mi>
                      <mml:mo>×</mml:mo>
                      <mml:mi>P</mml:mi>
                      <mml:mi>l</mml:mi>
                      <mml:mi>a</mml:mi>
                      <mml:mi>t</mml:mi>
                      <mml:mi>f</mml:mi>
                      <mml:mi>o</mml:mi>
                      <mml:mi>r</mml:mi>
                      <mml:mi>m</mml:mi>
                      <mml:mi>T</mml:mi>
                      <mml:mi>y</mml:mi>
                      <mml:mi>p</mml:mi>
                      <mml:mi>e</mml:mi>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>+</mml:mo>
                  <mml:mi>γ</mml:mi>
                  <mml:mo>×</mml:mo>
                  <mml:mi>C</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>n</mml:mi>
                  <mml:mi>t</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>l</mml:mi>
                  <mml:mi>V</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>i</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>b</mml:mi>
                  <mml:mi>l</mml:mi>
                  <mml:mi>e</mml:mi>
                  <mml:mi>s</mml:mi>
                  <mml:mo>+</mml:mo>
                  <mml:mi>ε</mml:mi>
                </mml:mtd>
              </mml:mtr>
            </mml:mtable>
          </mml:math>
        </disp-formula>
        <p>where: Control variables include text length, posting time (quarterly dummy variables), and engagement; Language Type (0 = English, 1 = Chinese); Platform Type (0 = community-led, 1 = algorithm-led).</p>
        <p>As shown in <bold>Table 5</bold>, the regression results of the model show that all core varia-bles exhibit significant characteristics, with a good model fit, which can effective-ly explain the generation mechanism of highly empowering content.</p>
        <p>Simple effect analysis shows that in both community-based and algorithm-based platforms, the proportion of highly empowering content for English users (38.6%, 19.7%) is higher than that for Chinese users (22.3%, 8.4%), and community-led platforms can alleviate the cross-language empowerment gap. The results of <bold>Table 5</bold> further confirm that after controlling relevant variables, language type still has a significant positive impact on the production of highly empowering content (<italic>β</italic> = 0.89, p &lt; 0.001). The interaction diagram (<xref ref-type="fig" rid="fig4">Figure 4</xref>) visually presents the marginal effects of the logistic regression model: the positive impact of English language on highly empowering content is more pronounced on community-led platforms, while algorithmic platforms amplify the cross-language empowerment gap. This visualization supports H3, confirming that language is an independent factor of the third financial digital divide, and platform type moderates the strength of this impact.</p>
        <fig id="fig4">
          <label>Figure 4</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId18.jpeg?20260228085951" />
        </fig>
        <p>Figure 4. Language and platform interaction effect diagram.</p>
        <p>Table 5. Logistic regression results.</p>
        <table-wrap id="tbl5">
          <label>Table 5</label>
          <table>
            <tbody>
              <tr>
                <td>Variables</td>
                <td>Coefficient</td>
                <td>Std. Error</td>
                <td>z-value</td>
                <td>p-value</td>
              </tr>
              <tr>
                <td>Intercept</td>
                <td>0.62</td>
                <td>0.08</td>
                <td>7.75</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Language Type (Chinese = 1)</td>
                <td>−0.95</td>
                <td>0.11</td>
                <td>−8.64</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Platform Type (Algorithmic = 1)</td>
                <td>−0.78</td>
                <td>0.10</td>
                <td>−7.80</td>
                <td>&lt;0.001</td>
              </tr>
              <tr>
                <td>Language × Platform Interaction</td>
                <td>−0.35</td>
                <td>0.13</td>
                <td>−2.69</td>
                <td>0.007</td>
              </tr>
              <tr>
                <td>Text Length</td>
                <td>0.02</td>
                <td>0.01</td>
                <td>2.01</td>
                <td>0.044</td>
              </tr>
              <tr>
                <td>Engagement Proxy Variable</td>
                <td>0.15</td>
                <td>0.06</td>
                <td>2.50</td>
                <td>0.012</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Model fit: McFadden pseudo R<sup>2</sup> = 0.32, AIC = 3862.5. Results show that language type, platform type, and their interaction all have significant negative impacts on highly empowering content, consistent with the original ANOVA conclusions but more compatible with the binary annotation data characteristics.</p>
      </sec>
      <sec id="sec4dot3">
        <title>4.3. Visual Analysis of Interaction Effects</title>
        <p>In order to intuitively present the interaction between language and platform and variable characteristics, this study constructed a dual verification system of “numerical testing + graphical evidence”. The grouped histogram (<xref ref-type="fig" rid="fig5">Figure 5</xref>) shows that the proportion of highly empowering content in the four groups of scenarios shows significant hierarchical differences, with the English-community group being the highest (38.6%) and the Chinese-algorithm group being the lowest (8.4%), confirming the empowering and compensatory role of community-led platforms for Chinese users. The core variable correlation heat map (<xref ref-type="fig" rid="fig6">Figure 6</xref>) shows that financial terminology coverage is positively correlated with the content professionalism score (r = 0.67, p &lt; 0.001) and negatively correlated with the semantic uncertainty score (r = −0.59, p &lt; 0.001), highlighting the central role of terminology standardization in content quality.</p>
        <fig id="fig5">
          <label>Figure 5</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId19.jpeg?20260228085951" />
        </fig>
        <p>Figure 5. Histogram of the proportion of highly empowering content in four groups of scenes.</p>
        <fig id="fig6">
          <label>Figure 6</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId20.jpeg?20260228085951" />
        </fig>
        <p>Figure 6. Core variable correlation heat map.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>5. Interpretation of Research Results and Literature Dialogue</title>
      <sec id="sec5dot1">
        <title>5.1. Interpretation of Core Research Results</title>
        <p>This study confirms that there is a significant financial expression ability gap among Chinese users, which is consistent with the research conclusions of Pan <italic>et al.</italic> (2020), and is visually verified by <xref ref-type="fig" rid="fig1">Figure 1</xref> (kernel density map). This gap is not caused by translation barriers, but by the superposition of multiple factors: the English financial terminology system is accurate and has high penetration rate, while Chinese terminology is mostly derived from translation, has ambiguous semantics and is not popular enough (as evidenced by the word cloud diagram in <xref ref-type="fig" rid="fig2">Figure 2</xref>); English communication focuses on logical data, while Chinese is biased towards experience and emotional expression; the English financial science popularization system is fragmented in Chinese. The difference in semantic uncertainty scores shows that Chinese users have shortcomings in term density and semantic accuracy. This indicator is quantified by the FinBERT model and fills the gap in cross-language scene measurement.</p>
        <p>5.1.1. The Impact Mechanism of Platform Affordances on Information Quality (the Reshaping of Content Professionalism by Platform Affordances)</p>
        <p>The test results of H2 confirm the interactive logic of media richness theory and platform affordance theory, that is, platform type affects the quality of financial content by adjusting the space for media richness. This mechanism is fully consistent with the theoretical analysis framework in Section 3.2 above. Through data mining on different platforms, this study found that platform mechanisms have a significant regulatory effect on information quality. After conducting multiple linear regression analysis, this study drew <xref ref-type="fig" rid="fig7">Figure 7</xref> (regression coefficient forest plot).</p>
        <fig id="fig7">
          <label>Figure 7</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId21.jpeg?20260228085953" />
        </fig>
        <p>Figure 7. Regression coefficient forest plot.</p>
        <p>As shown in <xref ref-type="fig" rid="fig7">Figure 7</xref>, the regression coefficient of platform type (Platform) is significantly positive, and the confidence interval does not cross the zero line, which statistically confirms that community-led platforms positively drive content professionalism. Combined with <xref ref-type="fig" rid="fig3">Figure 3</xref> (box plot of platform differences), it can be seen that the median score of community-led platforms (such as LinkedIn) is significantly higher than that of algorithm-based platforms (such as Weibo), which shows that “community affordances” have significantly improved the professional threshold of financial science popularization through its in-depth discussion mechanism.</p>
        <p>Algorithmic platforms are dominated by visibility affordances (recommendation streams account for &gt;70%), and circulate highly emotional, low-professional content through “emotional preferences-algorithmic push-interactive enhancement”. Most of the content is short videos and short texts, with low media richness. Community-led platforms are dominated by interactive affordances (attention streams account for &gt;60%). Users can independently filter information. The richness of media such as long texts and in-depth discussions is high, which is conducive to the dissemination of professional content. This result echoes relevant research views and adds to the new finding that “platform type can adjust the degree of inequality”. Community-led platforms can alleviate the information quality gap, and the conclusion is supported by visual results.</p>
        <p>5.1.2. The Interactive Effects and Theoretical Implications of Language and Platform</p>
        <p>The test results of H3 reveal the synergistic mechanism between language and platform, that is, the type of platform will moderate the impact of language differences on cognitive empowerment. This finding deepens the existing research’s understanding of the formation mechanism of the financial digital divide and improves the two-dimensional analysis framework proposed above.</p>
      </sec>
      <sec id="sec5dot2">
        <title>5.2. Hypothesis Testing and In-Depth Analysis of Interaction Effects (Dialogue with Existing Literature)</title>
        <p>In order to further explore the coupling relationship between language and platform, this study introduced interaction terms for testing. The results are shown in the visual presentation below.</p>
        <p>Observing <xref ref-type="fig" rid="fig8">Figure 8</xref>, we can see that the English text polyline shows a steeper upward trend on community-led platforms. This “non-parallel” relationship reveals the interactive effect of language and platform: the strong interactive affordances of community-led platforms can help Chinese users make up for terminology and semantic shortcomings and narrow the empowerment gap; algorithm-based platforms amplify this gap. Logistic regression shows that after controlling for platform type, language type still has a significant independent impact on the output of high-enabling content, indicating that language is the core factor of the third financial digital divide. This confirms that the cross-linguistic financial digital divide is the result of the combined effect of differences in language systems and heterogeneity of platform mechanisms, and provides empirical support for the construction of relevant theoretical frameworks.</p>
        <fig id="fig8">
          <label>Figure 8</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId22.jpeg?20260228085954" />
        </fig>
        <p>Figure 8. Language and platform interaction effect diagram.</p>
      </sec>
      <sec id="sec5dot3">
        <title>5.3. Theoretical Extension and Supplement</title>
        <p>This study improves the theoretical framework through dialogue with existing literature in three aspects: First, it expands the cross-linguistic application scenarios of the digital divide theory, confirms the existence of language heterogeneity in the second and third divides, and fills the theoretical gap; second, deepens the integration of the two major theories, reveals the interactive tension that high affordance does not equal high richness, and provides a new perspective for the application of traditional theories in the digital age; third, extends the social science boundaries of the FinBERT model, combines it with the financial digital divide theory, quantifies financial literacy and empowerment effects, and improves the methodology.</p>
        <p>At the level of NLP technology implementation, this study overcomes the multiple challenges of FinBERT cross-language fine-tuning, including differences in Chinese and English word segmentation, the impact of unregistered words, cross-language embedding alignment errors, etc., which not only improves measurement accuracy, but also provides empirical parameters for fine-tuning cross-language financial models.</p>
      </sec>
      <sec id="sec5dot4">
        <title>5.4. Research Limitations</title>
        <p>Although this study has made breakthroughs in theory and method, it still has three limitations in combination with the previous research design, which points out the direction for subsequent research: First, at the data level, the sample only covers the entire year of 2024, lacking longitudinal tracking, making it difficult to capture the dynamic evolution characteristics of the financial digital divide; at the same time, the data only comes from 10 mainstream platforms, and does not cover niche financial communities and regional platforms. There are certain limitations in sample representativeness, which is directly related to the definition of the data collection scope in Section 3.3 above. In addition, although the use of 50 tokens standardization improves cross-language comparability, it still cannot completely eliminate the inherent differences in Chinese and English language structures (such as no inflection in Chinese and different subword splitting logic in English), which may have a slight impact on the absolute comparison of indicators such as term coverage. Subsequent optimization can be further combined with semantic embedding vectors.</p>
        <p>Second, at the method level, although cross-language embedding alignment is achieved through unified fine-tuning and the MUSE tool, there are still potential deficiencies in the absolute comparability of the Chinese and English FinBERT models, and the quantification of semantic ambiguity may be affected by language and cultural differences. In addition, although propensity score matching controls some interference variables, it still cannot completely eliminate endogeneity problems. The credibility of causal inference needs to be further improved, and there is a certain subjectivity in the semantic annotation of some texts, which may affect the analysis accuracy. This is also a potential limitation mentioned in the research method section of Section 3.4.</p>
        <p>Third, at the variable level, individual user characteristics (such as education level and financial experience) are not included in this study, making it difficult for this study to reveal the moderating effect of individual heterogeneity on the financial digital divide. At the same time, although the definition of high-enabling content refers to the existing framework, it still has a certain degree of subjectivity and may affect measurement accuracy. This limitation also provides a direction for the improvement of the variable system in subsequent research.</p>
      </sec>
    </sec>
    <sec id="sec6">
      <title>6. Conclusions and Future Research Directions</title>
      <sec id="sec6dot1">
        <title>6.1. Research Conclusions</title>
        <p>This study is supported by the digital divide theory, media richness theory and platform affordance perspective. Based on the research scope and data collection plan defined in Section 3.3 above, this study selects 35,000 financial text data from 10 Chinese and English social media platforms. It uses the unified and fine-tuned bilingual FinBERT model to combine statistical testing and visual analysis methods to systematically test the impact of language type and platform type on the financial digital divide and the interactive effect of the two. Echoing the three major research hypotheses proposed above, the following core conclusions are drawn:</p>
        <p>This study reveals the logic of generating the quality of financial science popularization in the social media environment through multi-dimensional empirical analysis of 35,000 pieces of data. In order to systematically present the research findings, this study constructed <xref ref-type="fig" rid="fig9">Figure 9</xref> (social media financial information quality impact mechanism model).</p>
        <fig id="fig9">
          <label>Figure 9</label>
          <graphic xlink:href="https://html.scirp.org/file/1733460-rId23.jpeg?20260228085957" />
        </fig>
        <p>Figure 9. Social media financial information quality impact mechanism model.</p>
        <p>Summary of conclusion: This research model reveals that financial information quality is driven by both “language barriers” and “platform logic”. The current situation of low empowerment of Chinese financial information is not irreversible and can be alleviated by optimizing the platform recommendation logic and introducing a standardized terminology system. The core conclusions are as follows: First, there is a significant financial expression ability gap in cross-language contexts. The structural differences in the Chinese and English semantic expression systems show that Chinese users have lower coverage of financial terms and higher semantic uncertainty. After controlling for observable characteristics such as text length, Chinese and English users show systematic differences in financial expression ability. This difference is highly consistent with the theoretical expectations of the second financial digital divide. Second, differences in platform affordances create an information quality gap. Algorithmic platforms are more likely to disseminate highly emotional and low-professional content, while community-led platforms are more conducive to the dissemination of professional content. This difference exists across languages and will exacerbate the second gap. Third, the interaction between language and platform significantly affects the cognitive empowerment effect. Language itself is an independent factor in the third gap. Community-led platforms can alleviate the cross-language empowerment gap, while algorithmic platforms amplify the gap, completing the “language-platform” two-dimensional analysis framework.</p>
      </sec>
      <sec id="sec6dot2">
        <title>6.2. Practical Implications</title>
        <p>Platform operators need to optimize algorithms, introduce collaborative annotation and professional review mechanisms, increase the weight of standardized financial terminology content, reduce emotional content push, and increase in-depth content recommendations; community-led platforms strengthen interaction and content screening, and optimize term semantic calibration across language platforms.</p>
        <p>Financial science popularization workers should build a systematic Chinese financial terminology system to reduce semantic ambiguity, create science popularization content that combines fragmentation and systematization, implement differentiated science popularization for different platforms, and strengthen terminology application scenario training.</p>
        <p>Policymakers need to incorporate the cross-language financial digital divide into inclusive financial policies, increase support for Chinese financial science popularization, standardize platform content distribution, promote the standardization of cross-border financial information services, and formulate differentiated regulatory policies.</p>
      </sec>
      <sec id="sec6dot3">
        <title>6.3. Future Prospects</title>
        <p>In the future, research can be deepened from three aspects: first, optimize data and methods, use longitudinal tracking data, expand sample coverage, and combine causal inference methods to reduce endogeneity; second, expand theories and perspectives, explore the moderating role of individual user characteristics, integrate cross-cultural communication theory, and verify the cross-language universality of conclusions; third, extend practical applications, develop financial literacy improvement tools and algorithm optimization plans, conduct cross-regional comparative studies, and combine large language model technology to develop cross-language financial content conversion tools to narrow the cross-language financial digital divide.</p>
      </sec>
    </sec>
    <sec id="sec7">
      <title>Appendix A: Examples of Highly Empowering/Non-Empowering Text Annotation</title>
      <p>Table A1. Comparison table of examples of highly empowered and non-empowered text annotations.</p>
      <table-wrap id="tbl6">
        <label>Table 6</label>
        <table>
          <tbody>
            <tr>
              <td>language type</td>
              <td>text type</td>
              <td>original text</td>
              <td>Reason for labeling</td>
            </tr>
            <tr>
              <td>Chinese</td>
              <td>Highly empowering</td>
              <td>The Fed’s interest rate hike will increase bond yields, and it is recommended to reduce duration allocation to avoid interest rate risks. The current 10-year government bond yield has exceeded 3.2%, and we need to focus on policy changes in the short term.</td>
              <td>It includes mechanism explanation (interest rate increase → rising yield), risk warning (avoiding interest rate risk), data support (10-year government bond yield is 3.2%) and decision-making reference (reducing duration allocation), which fully meets the three empowerment dimensions of “completeness of mechanism explanation, clarity of risk warning, and decision-making reference value”.</td>
            </tr>
            <tr>
              <td>Chinese</td>
              <td>non- empowering</td>
              <td>This fund is so profitable, everyone should buy it quickly. If it is too late, you will have no chance!</td>
              <td>It contains only emotional recommendations without any financial logic explanation, risk warning or data support, and does not meet the definition of empowering content.</td>
            </tr>
            <tr>
              <td>Chinese</td>
              <td>Highly empowering</td>
              <td>During the A-share annual reporting season, we need to focus on the matching between net profit growth and cash flow. If net profit increases but cash flow from operating activities is negative, there may be a revenue recognition bias. It is recommended to further verify it based on the accounts receivable turnover rate.</td>
              <td>It includes interpretation of core indicators (net profit growth, cash flow, accounts receivable turnover rate), risk warning (revenue recognition deviation) and verification methods, and has strong cognitive empowerment value.</td>
            </tr>
            <tr>
              <td>Chinese</td>
              <td>non- empowering</td>
              <td>You are right to follow me and buy stocks. What you bought yesterday has gone up 5 points today!</td>
              <td>It only shares personal investment returns, without professional analysis and risk warnings, and is an empirical and emotional expression.</td>
            </tr>
            <tr>
              <td>English</td>
              <td>Highly empowering</td>
              <td>Fed rate hikes increase bond yields; investors should reduce duration to mitigate interest rate risk. The 10-year Treasury yield has exceeded 3.2%, so short-term focus should be on policy changes.</td>
              <td>It includes mechanism explanation, risk warning, data support and decision-making suggestions. It meets all three empowerment dimensions and meets the high-enabling content standards.</td>
            </tr>
            <tr>
              <td>English</td>
              <td>non- empowering</td>
              <td>This fund is making a killing—hurry up and buy before it’s too late!</td>
              <td>It only contains emotional and inflammatory expressions, without professional financial analysis, risk warnings or logical support, and does not constitute empowering content.</td>
            </tr>
            <tr>
              <td>English</td>
              <td>Highly empowering</td>
              <td>During the A-share annual report season, focus on the matching degree between net profit growth and cash flow. Negative operating cash flow with rising net profit may indicate revenue recognition deviations; verify with accounts receivable turnover.</td>
              <td>Covers indicator interpretation, risk warning and verification methods, and has complete cognitive empowerment logic.</td>
            </tr>
            <tr>
              <td>English</td>
              <td>non- empowering</td>
              <td>Follow my stock picks and you’ll never lose—I bought this yesterday and it’s already up 5% today!</td>
              <td>It only shares short-term gains, without professional analysis and risk warning, and is an emotional expression based on experience.</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <p>1) Qualifications of annotators: The annotation team for this study is a bilingual financial field expert with 10 years of working experience in the information department of China’s state-owned banks and 10 years of experience in financial terminology training at Citibank Singapore Branch to ensure the professionalism and consistency of annotation standards. 2) Annotation consistency test: The annotation consistency is verified through Cohen’s Kappa test. The Kappa value is 0.82, which meets the annotation quality standard for quantitative research (Kappa ≥ 0.8 is highly consistent). 3) Definition of empowerment dimensions: Highly empowering content must simultaneously meet the three dimensions of “completeness of mechanism explanation”, “clarity of risk warning” and “decision-making reference value”. If one of them is missing, it will be judged as non-enabling content.</p>
    </sec>
    <sec id="sec8">
      <title>Appendix B: Sentiment Score Distribution and Threshold Calibration</title>
      <p>Table B1. Sentiment score percentiles of annotated texts.</p>
      <table-wrap id="tbl7">
        <label>Table 7</label>
        <table>
          <tbody>
            <tr>
              <td>Percentile</td>
              <td>Sentiment Score</td>
              <td>Interpretation</td>
            </tr>
            <tr>
              <td>5th</td>
              <td>−0.82</td>
              <td>Lower bound of extreme negative sentiment</td>
            </tr>
            <tr>
              <td>10th</td>
              <td>−0.70</td>
              <td>Selected threshold for negative polarization</td>
            </tr>
            <tr>
              <td>90th</td>
              <td>0.70</td>
              <td>Selected threshold for positive polarization</td>
            </tr>
            <tr>
              <td>95th</td>
              <td>0.83</td>
              <td>Upper bound of extreme positive sentiment</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <fig id="fig10">
        <label>Figure 10</label>
        <graphic xlink:href="https://html.scirp.org/file/1733460-rId32.jpeg?20260228090000" />
      </fig>
      <p>Note: Insert histogram with x-axis = sentiment score (−1 - 1), y-axis = frequency; mark thresholds at ±0.7.</p>
      <p>Figure B1. Histogram of sentiment score distribution.</p>
    </sec>
    <sec id="sec9">
      <title>Appendix C: PSM Matching Balance Diagnostic Plots</title>
      <fig id="fig11">
        <label>Figure 11</label>
        <graphic xlink:href="https://html.scirp.org/file/1733460-rId33.jpeg?20260228090000" />
      </fig>
      <p>Note: Insert horizontal bar plot comparing SMD values of all covariates before and after matching.</p>
      <p>Figure C1. Standardized Mean Difference (SMD) plot before and after matching.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Gupta, S. and Chen, H. (2020) Social Media and Financial Information Ecosystem: A Systematic Review. <italic>Journal of Business Ethics</italic>, 165, 457-478.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Gupta, S.</string-name>
              <string-name>Chen, H.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Social Media and Financial Information Ecosystem: A Systematic Review</article-title>
            <source>Journal of Business Ethics</source>
            <volume>165</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Chu, Y., Li, M. and Wang, H. (2022) A Comparative Study of Financial Information Sharing on Chinese and English Social Media Platforms. <italic>Journal of Computer</italic>- <italic>Mediated Communication</italic>, 27, 892-910.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Chu, Y.</string-name>
              <string-name>Li, M.</string-name>
              <string-name>Wang, H.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>A Comparative Study of Financial Information Sharing on Chinese and English Social Media Platforms</article-title>
            <source>Journal of Computer-Mediated Communication</source>
            <volume>27</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Van Dijk, J. (2006) The Deepening Divide: Inequality in the Information Society. Sage Publications. https://doi.org/10.4135/9781452229812 <pub-id pub-id-type="doi">10.4135/9781452229812</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4135/9781452229812">https://doi.org/10.4135/9781452229812</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Dijk, J.</string-name>
            </person-group>
            <year>2006</year>
            <article-title>The Deepening Divide: Inequality in the Information Society</article-title>
            <pub-id pub-id-type="doi">10.4135/9781452229812</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Yang, C., Zhang, Y. and Liu, J. (2020) Cross-Lingual Adaptability of FinBERT Model in Financial Text Analysis. <italic>IEEE Transactions on Knowledge and Data Engineering</italic>, 34, 3892-3905.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Yang, C.</string-name>
              <string-name>Zhang, Y.</string-name>
              <string-name>Liu, J.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Cross-Lingual Adaptability of FinBERT Model in Financial Text Analysis</article-title>
            <source>IEEE Transactions on Knowledge and Data Engineering</source>
            <volume>34</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Husin, N. (2024) Social Media Content and Youth Financial Decision-Making: A Quantitative Analysis. <italic>Journal of Youth Studies</italic>, 28, 213-235.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Husin, N.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Social Media Content and Youth Financial Decision-Making: A Quantitative Analysis</article-title>
            <source>Journal of Youth Studies</source>
            <volume>28</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Hofstede, G. (2011) Dimensionalizing Cultures: The Hofstede Model in Context. <italic>Online Readings in Psychology and Culture</italic>, 2, 1-26. https://doi.org/10.9707/2307-0919.1014 <pub-id pub-id-type="doi">10.9707/2307-0919.1014</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.9707/2307-0919.1014">https://doi.org/10.9707/2307-0919.1014</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Hofstede, G.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Dimensionalizing Cultures: The Hofstede Model in Context</article-title>
            <source>Online Readings in Psychology and Culture</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.9707/2307-0919.1014</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Stolper, O.A. and Walter, A. (2017) Financial Literacy, Financial Advice, and Financial Behavior. <italic>Journal of Business Economics</italic>, 87, 581-643. https://doi.org/10.1007/s11573-017-0853-9 <pub-id pub-id-type="doi">10.1007/s11573-017-0853-9</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s11573-017-0853-9">https://doi.org/10.1007/s11573-017-0853-9</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Stolper, O.A.</string-name>
              <string-name>Walter, A.</string-name>
              <string-name>Literacy, F</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Financial Literacy, Financial Advice, and Financial Behavior</article-title>
            <source>Journal of Business Economics</source>
            <volume>87</volume>
            <pub-id pub-id-type="doi">10.1007/s11573-017-0853-9</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Vassilakopoulou, A. and Hustad, J. (2023) Algorithmic Bias and Financial Digital Divide in the Social Media Era. <italic>New Media &amp; Society</italic>, 25, 1456-1478.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Vassilakopoulou, A.</string-name>
              <string-name>Hustad, J.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Algorithmic Bias and Financial Digital Divide in the Social Media Era</article-title>
            <source>New Media &amp; Society</source>
            <volume>25</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Li, Z. and Wang, L. (2023) Integrating NLP Technology with Social Science Theories: A Case Study of Financial Digital Inequality Research. <italic>Social Science Computer Review</italic>, 41, 789-806.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Li, Z.</string-name>
              <string-name>Wang, L.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Integrating NLP Technology with Social Science Theories: A Case Study of Financial Digital Inequality Research</article-title>
            <source>Social Science Computer Review</source>
            <volume>41</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Daft, R.L. and Lengel, R.H. (1986) Organizational Information Requirements, Media Richness and Structural Design. <italic>Management Science</italic>, 32, 554-571. https://doi.org/10.1287/mnsc.32.5.554 <pub-id pub-id-type="doi">10.1287/mnsc.32.5.554</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1287/mnsc.32.5.554">https://doi.org/10.1287/mnsc.32.5.554</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Daft, R.L.</string-name>
              <string-name>Lengel, R.H.</string-name>
              <string-name>Requirements, M</string-name>
            </person-group>
            <year>1986</year>
            <article-title>Organizational Information Requirements, Media Richness and Structural Design</article-title>
            <source>Management Science</source>
            <volume>32</volume>
            <pub-id pub-id-type="doi">10.1287/mnsc.32.5.554</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kim, J., Lee, S. and Park, H. (2021) Cross-Platform Differences in Information Dissemination: A Perspective of Media Richness Theory. <italic>Journal of Communication</italic>, 71, 289-312.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kim, J.</string-name>
              <string-name>Lee, S.</string-name>
              <string-name>Park, H.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Cross-Platform Differences in Information Dissemination: A Perspective of Media Richness Theory</article-title>
            <source>Journal of Communication</source>
            <volume>71</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Angelica, M., Rossi, F. and Verdi, C. (2023) The Impact of Social Media Financial Content on Financial Literacy Enhancement. <italic>Journal of Consumer Affairs</italic>, 57, 189-210.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Angelica, M.</string-name>
              <string-name>Rossi, F.</string-name>
              <string-name>Verdi, C.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>The Impact of Social Media Financial Content on Financial Literacy Enhancement</article-title>
            <source>Journal of Consumer Affairs</source>
            <volume>57</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Lu, Y., Chen, J. and Zhang, Q. (2023) Internet Infrastructure and Rural Financial Inclusion in China. <italic>China Economic Review</italic>, 79, Article 101892.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Lu, Y.</string-name>
              <string-name>Chen, J.</string-name>
              <string-name>Zhang, Q.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Internet Infrastructure and Rural Financial Inclusion in China</article-title>
            <source>China Economic Review</source>
            <volume>79</volume>
            <elocation-id>101892</elocation-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Lusardi, A. and Mitchell, O.S. (2014) Financial Literacy and Planning: Implications for Retirement Wellbeing. <italic>Journal of Economic Perspectives</italic>, 28, 43-60.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Lusardi, A.</string-name>
              <string-name>Mitchell, O.S.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Financial Literacy and Planning: Implications for Retirement Wellbeing</article-title>
            <source>Journal of Economic Perspectives</source>
            <volume>28</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Aissaoui, M. (2022) Objective Measurement of Financial Literacy Using Digital Behavioral Data. <italic>Journal of Financial Counseling and Planning</italic>, 33, 156-172.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Aissaoui, M.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Objective Measurement of Financial Literacy Using Digital Behavioral Data</article-title>
            <source>Journal of Financial Counseling and Planning</source>
            <volume>33</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Bucher, T. and Helmond, A. (2018) The Work of Platforms: Reflexivity and the New Media Event. <italic>Information</italic>, <italic>Communication &amp; Society</italic>, 21, 22-39.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Bucher, T.</string-name>
              <string-name>Helmond, A.</string-name>
              <string-name>Information, C</string-name>
            </person-group>
            <year>2018</year>
            <article-title>The Work of Platforms: Reflexivity and the New Media Event</article-title>
            <source>Information</source>
            <volume>21</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Burke, K. and Hung, M. (2021) User Engagement in Online Learning Communities: Integrating Media Richness and Affordance Theory. <italic>Computers &amp; Education</italic>, 175, Article 104389.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Burke, K.</string-name>
              <string-name>Hung, M.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>User Engagement in Online Learning Communities: Integrating Media Richness and Affordance Theory</article-title>
            <source>Computers &amp; Education</source>
            <volume>175</volume>
            <elocation-id>104389</elocation-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Kim, H., Park, J. and Lee, J. (2021) Research Gaps in Cross-Platform Information Behavior Studies. <italic>Library &amp; Information Science Research</italic>, 43, Article 101032.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Kim, H.</string-name>
              <string-name>Park, J.</string-name>
              <string-name>Lee, J.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Research Gaps in Cross-Platform Information Behavior Studies</article-title>
            <source>Library &amp; Information Science Research</source>
            <volume>43</volume>
            <elocation-id>101032</elocation-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Huang, Y., Wang, S. and Li, C. (2023) A Multidimensional Evaluation Framework for Financial Information Quality on Social Media. <italic>Journal of Management Information Systems</italic>, 40, 567-598.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Huang, Y.</string-name>
              <string-name>Wang, S.</string-name>
              <string-name>Li, C.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>A Multidimensional Evaluation Framework for Financial Information Quality on Social Media</article-title>
            <source>Journal of Management Information Systems</source>
            <volume>40</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Amaral, G. and Kolsarici, C. (2020) A Meta-Analysis of Financial Literacy Measurement Methods. <italic>Journal of Economic Surveys</italic>, 34, 890-912.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Amaral, G.</string-name>
              <string-name>Kolsarici, C.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>A Meta-Analysis of Financial Literacy Measurement Methods</article-title>
            <source>Journal of Economic Surveys</source>
            <volume>34</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Gordon, M.L., Lam, M.S., Park, J.S., Patel, K., Hancock, J., Hashimoto, T., <italic>et al</italic>. (2022) Jury Learning: Integrating Dissenting Voices into Machine Learning Models. In: <italic>Proceedings of the</italic>2022 <italic>CHI Conference on Human Factors in Computing Systems</italic>( <italic>CHI</italic>’22), Article No.115, 1-19. https://doi.org/10.1145/3491102.3502004 <pub-id pub-id-type="doi">10.1145/3491102.3502004</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1145/3491102.3502004">https://doi.org/10.1145/3491102.3502004</ext-link></mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Gordon, M.L.</string-name>
              <string-name>Lam, M.S.</string-name>
              <string-name>Park, J.S.</string-name>
              <string-name>Patel, K.</string-name>
              <string-name>Hancock, J.</string-name>
              <string-name>Hashimoto, T.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Jury Learning: Integrating Dissenting Voices into Machine Learning Models</article-title>
            <source>In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ’22)</source>
            <volume>1</volume>
            <elocation-id>No.115</elocation-id>
            <pub-id pub-id-type="doi">10.1145/3491102.3502004</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Araci, D. (2019) FinBERT: Financial Sentiment Analysis with Pre-Trained Language Models. http://arxiv.org/abs/1908.10063</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Araci, D.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>FinBERT: Financial Sentiment Analysis with Pre-Trained Language Models</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Huang, A.H., Wang, H. and Yang, Y. (2023) FinBERT: A Large Language Model for Extracting Information from Financial Text. <italic>Contemporary Accounting Research</italic>, 40, 806-841. https://doi.org/10.1111/1911-3846.12832 <pub-id pub-id-type="doi">10.1111/1911-3846.12832</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1111/1911-3846.12832">https://doi.org/10.1111/1911-3846.12832</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Huang, A.H.</string-name>
              <string-name>Wang, H.</string-name>
              <string-name>Yang, Y.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>FinBERT: A Large Language Model for Extracting Information from Financial Text</article-title>
            <source>Contemporary Accounting Research</source>
            <volume>40</volume>
            <pub-id pub-id-type="doi">10.1111/1911-3846.12832</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Liu, Y., Chen, X. and Wang, Z. (2020) Comparative Analysis of FinBERT and General BERT in Financial Text Processing. <italic>Journal of Financial Data Science</italic>, 2, 78-92.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Liu, Y.</string-name>
              <string-name>Chen, X.</string-name>
              <string-name>Wang, Z.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Comparative Analysis of FinBERT and General BERT in Financial Text Processing</article-title>
            <source>Journal of Financial Data Science</source>
            <volume>2</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Fjellstrom, M. (2022) Stock Market Volatility Prediction Using FinBERT and LSTM Model. <italic>Journal of Forecasting</italic>, 41, 987-1002.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Fjellstrom, M.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Stock Market Volatility Prediction Using FinBERT and LSTM Model</article-title>
            <source>Journal of Forecasting</source>
            <volume>41</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="confproc">Chin, T., Lee, K. and Ng, W. (2024) Extracting Investment Signals from Earnings Conference Calls Using FinBERT. <italic>Accounting Horizons</italic>, 38, 45-62.</mixed-citation>
          <element-citation publication-type="confproc">
            <person-group person-group-type="author">
              <string-name>Chin, T.</string-name>
              <string-name>Lee, K.</string-name>
              <string-name>Ng, W.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Extracting Investment Signals from Earnings Conference Calls Using FinBERT</article-title>
            <source>Accounting Horizons</source>
            <volume>38</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Saunders, M., Lewis, P. and Thornhill, A. (2019) Research Methods for Business Students. 8th Edition, Pearson Education Limited.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Saunders, M.</string-name>
              <string-name>Lewis, P.</string-name>
              <string-name>Thornhill, A.</string-name>
              <string-name>Edition, P</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Research Methods for Business Students</article-title>
            <source>8th Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Hair, J.F., Black, W.C., Babin, B.J. and Anderson, R.E. (2014) Multivariate Data Analysis. 8th Edition, Pearson Prentice Hall.</mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Hair, J.F.</string-name>
              <string-name>Black, W.C.</string-name>
              <string-name>Babin, B.J.</string-name>
              <string-name>Anderson, R.E.</string-name>
              <string-name>Edition, P</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Multivariate Data Analysis</article-title>
            <source>8th Edition</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B29">
        <label>29.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Hansen, S., McMahon, M. and Prat, A. (2018) Transparency and Deliberation within the FOMC: A Computational Linguistics Approach. <italic>The Quarterly Journal of Economics</italic>, 133, 801-870. https://doi.org/10.1093/qje/qjx045 <pub-id pub-id-type="doi">10.1093/qje/qjx045</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1093/qje/qjx045">https://doi.org/10.1093/qje/qjx045</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Hansen, S.</string-name>
              <string-name>McMahon, M.</string-name>
              <string-name>Prat, A.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Transparency and Deliberation within the FOMC: A Computational Linguistics Approach</article-title>
            <source>The Quarterly Journal of Economics</source>
            <volume>133</volume>
            <pub-id pub-id-type="doi">10.1093/qje/qjx045</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B30">
        <label>30.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Pan, L., <italic>et al</italic>. (2020) Cross-Cultural Differences in Financial Discourse: A Corpus-Based Study. <italic>Journal of Pragmatics</italic>, 155, 120-135.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Pan, L.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Cross-Cultural Differences in Financial Discourse: A Corpus-Based Study</article-title>
            <source>Journal of Pragmatics</source>
            <volume>155</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>