<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jdaip</journal-id>
      <journal-title-group>
        <journal-title>Journal of Data Analysis and Information Processing</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-7203</issn>
      <issn pub-type="ppub">2327-7211</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jdaip.2026.142012</article-id>
      <article-id pub-id-type="publisher-id">jdaip-151201</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Full-Fidelity Semantic Aggregation: Navigating Datasets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <contrib-id contrib-id-type="orcid">0009-0004-0938-1783</contrib-id>
          <name name-style="western">
            <surname>Mallgren</surname>
            <given-names>Anthony Brian</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Independent Researcher, New York, NY, USA </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declare no conflicts of interest regarding the publication of this paper.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>01</day>
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>05</month>
        <year>2026</year>
      </pub-date>
      <volume>14</volume>
      <issue>02</issue>
      <fpage>247</fpage>
      <lpage>253</lpage>
      <history>
        <date date-type="received">
          <day>13</day>
          <month>03</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>09</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="published">
          <day>12</day>
          <month>05</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jdaip.2026.142012">https://doi.org/10.4236/jdaip.2026.142012</self-uri>
      <abstract>
        <p>Full-fidelity semantic aggregation presents significant advantages over utilizing publicly facing contemporary LLMs hosted by other organizations (e.g., cost, fewer required resources, interoperability, tail analysis, et cetera). Full-fidelity models can be trained with low-end/legacy hardware, become functional with virtually any dataset, and can be used in cross-dataset analysis. The model of truth approach that is generally made available to the general public seems to target a general audience by means of prescriptive heuristics, which involve issues related to dogma. It also lacks a more holistic familiarity through breadth in data. The approach outlined here exposes what the equivalent of neural network weights, or as it is labeled herein, conduciveness. It allows additional flexibility in assessments of heavily opinionated subjects. It allows analysis of divergence in variance and intensity. This method also opens a wide scope of innovation to occur in the academic world. The intent here is to provide a brief evaluation to prove the viability of the method.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Data Analysis</kwd>
        <kwd>Computational Semantics</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Unstructured Aggregation</kwd>
        <kwd>Full-Fidelity Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Some technologists are focusing on utilizing models to produce what they term as artificial intelligence [<xref ref-type="bibr" rid="B1">1</xref>]. One method being utilized in an attempt to accomplish this is training large language models (LLMs). This requires substantial resources [<xref ref-type="bibr" rid="B2">2</xref>], which makes them heavily dependent upon what has thus far been corporate sponsorship. Additionally, this approach to data analysis consumes a significant amount of energy [<xref ref-type="bibr" rid="B3">3</xref>]. Some activity centers around constraining these models to make them perform as desired [<xref ref-type="bibr" rid="B4">4</xref>]. There are articles covering how these models may be utilized in data analysis [<xref ref-type="bibr" rid="B5">5</xref>].</p>
      <p>There are a few assumptions that are being made. One, that data can be used primarily in a humanistic aspect, <italic>i.e.</italic>, that the primary goal is developmentally human centric. Two, that energy should be utilized as efficiently as possible to achieve this goal. And three, that democratizing data analysis of larger datasets, addressed more specifically herein, of unstructured datasets, is generally agreeable within and for society. These are the justifications behind the work presented here. Some advantages to the approach presented are that larger datasets can be wrangled and navigated with fewer resources, a more holistic understanding of a dataset is allowed, and dogmatic opinions formed by artificial intelligence companies are avoided.</p>
      <p>What is presented herein is semantic aggregation. The examples presented are of linguistic datasets, though they could be utilized for other datasets.</p>
    </sec>
    <sec id="sec2">
      <title>2. Methods</title>
      <sec id="sec2dot1">
        <title>2.1. Development</title>
        <p>The resources required to set up a semantic aggregation data analysis environment are relatively low cost, or even free. In the most extreme data savings scenarios, a free environment could be used, such as those provided through <ext-link ext-link-type="uri" xlink:href="https://ifastnet.com">https://ifastnet.com</ext-link>, or through provider partners such as <ext-link ext-link-type="uri" xlink:href="https://infinityfree.com">https://infinityfree.com</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://aeonfree.com">https://aeonfree.com</ext-link>, etc. This allows free databases and web hosting. Due to some of the security restrictions and limited functionality in performing testing, a JavaScript agent was created to parse and process data. The intent of a security mechanism enforced in utilizing these services is that the web request is not cross-domain and is being executed in a browser which executes JavaScript. It is worth noting that batching is useful here, as the number of transactions is throttled. Also, occasionally, it may be necessary to file a ticket for review, as there are automated mechanisms that shut down services. If one can afford a few dollars a month (&lt;$10), this opens access to providers that allow the use of expanded server-side execution, such as CRON jobs. This enables 24/7/365 processing and virtually unlimited sized datasets. If there is a larger budget, cloud providers can be utilized. Or, if private data suffices, and one has access to their own equipment, a local environment may be set up and utilized in developing aggregated maps.</p>
        <p>Once an environment is settled upon, there is one method that has been found to be efficient in analyzing larger datasets; this can be termed the step-down method. A first table could simply keep the word count of the words occurring in a dataset and keep the count of occurrences in that dataset. For example, as partially shown in <bold>Table 1</bold>:</p>
        <p><bold>Table 1</bold><bold>.</bold> Word occurrence count (conduciveness).</p>
        <table-wrap id="tbl1">
          <label>Table 1</label>
          <table>
            <tbody>
              <tr>
                <td>Word</td>
                <td>Conduciveness</td>
              </tr>
              <tr>
                <td>the</td>
                <td>2,047,163</td>
              </tr>
              <tr>
                <td>of</td>
                <td>1,383,152</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The next step-down table, as partially shown in <bold>Table 2</bold>:</p>
        <p><bold>Table 2</bold><bold>.</bold> First expanded extension (may be dynamically named, e.g., “the”).</p>
        <table-wrap id="tbl2">
          <label>Table 2</label>
          <table>
            <tbody>
              <tr>
                <td>Word</td>
                <td>Following Word</td>
                <td>Conduciveness</td>
                <td>Stepdown Table</td>
              </tr>
              <tr>
                <td>the</td>
                <td>Secretary</td>
                <td>125,050</td>
                <td>the1</td>
              </tr>
              <tr>
                <td>the</td>
                <td>United</td>
                <td>78,111</td>
                <td>the2</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>And the next step-down table would be, as partially shown in <bold>Table 3</bold>:</p>
        <p><bold>Table 3</bold><bold>.</bold> Second Expanded Extension (may be dynamically named with partition #, e.g., “the1”).</p>
        <table-wrap id="tbl3">
          <label>Table 3</label>
          <table>
            <tbody>
              <tr>
                <td>Word</td>
                <td>Following Word</td>
                <td>Second Following Word</td>
                <td>Conduciveness</td>
                <td>Stepdown Table</td>
              </tr>
              <tr>
                <td>the</td>
                <td>Secretary</td>
                <td>of</td>
                <td>41,234</td>
                <td>the_Secretary1</td>
              </tr>
              <tr>
                <td>the</td>
                <td>Secretary</td>
                <td>shall</td>
                <td>30,236</td>
                <td>the_Secretary1</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>The stepdown method allows efficient indexing, allows partitioning, et cetera. The exact allocation and optimization methods for data restructuring will depend on the size of the dataset, the goals, et cetera.</p>
        <p>This method allows a better understanding of larger datasets and enables aggregate navigation. A comparison method between datasets can be utilized. This can involve comparing heterogeneous datasets or time slices of homogeneous datasets. For example, a report may be, as partially shown in <bold>Table 4</bold>:</p>
        <p><bold>Table 4</bold><bold>.</bold>Dataset comparison result.</p>
        <table-wrap id="tbl4">
          <label>Table 4</label>
          <table>
            <tbody>
              <tr>
                <td>Word</td>
                <td>Following Word</td>
                <td>Position in Dataset 1</td>
                <td>Position in Dataset 2</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>of</td>
                <td>1</td>
                <td>1</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>or</td>
                <td>2</td>
                <td>4</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>and</td>
                <td>3</td>
                <td>0</td>
              </tr>
              <tr>
                <td>Governmental</td>
                <td>Affairs</td>
                <td>4</td>
                <td>2</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>to</td>
                <td>5</td>
                <td>5</td>
              </tr>
              <tr>
                <td>Governments</td>
                <td>and</td>
                <td>6</td>
                <td>6</td>
              </tr>
              <tr>
                <td>Governmental</td>
                <td>organizations</td>
                <td>7</td>
                <td>7</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>in</td>
                <td>8</td>
                <td>0</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>Act</td>
                <td>9</td>
                <td>10</td>
              </tr>
              <tr>
                <td>Government</td>
                <td>Accountability</td>
                <td>10</td>
                <td>8</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>Additional information can be added to the dataset. For example, parts of speech. This can be done with higher accuracy using a parts-of-speech tagger, or a dictionary approach can be used to label the possible parts of speech. This allows more targeted queries. Also, it allows heuristic approaches to data analysis. For example, word substitution analysis. For example, government [noun], or government [verb]. This can be expanded and distilled into larger models. Comparisons can also be made. For example, the President [verb] versus the Secretary [verb], and compared as previously done with heterogeneous or homogeneous datasets.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Analysis</title>
        <p>The analysis that was performed was simple. As previously mentioned, free platforms were assessed for their usefulness in getting started with such data analysis. It seems that anyone with access to a library may begin partaking in this type of analysis, though an active session is required to process data. Cell phones may be used to process data after authoring the scripts, but the tab generally goes inactive, and settings may need to be modified for continuous processing. As previously mentioned, requests are throttled, and as was implied, space is limited. Therefore, it may be seen as a way to get started and bide one’s time, though not a means to an end (unless the datasets are relatively small).</p>
        <p>The next level up, <italic>i.e.</italic>, low-priced shared hosting, which is typically affordable even on the lowest tier of government benefits, offers 24/7/365 processing possibilities and virtually unlimited-sized datasets. It is noteworthy that, as seems to be the nature with information technology systems, progression tracking and robust exception handling may make the experience more agreeable.</p>
        <p>Cloud providers were not utilized in the analysis, though the same approach for progression tracking and exception handling seems relevant, as other systems seem to err occasionally during data collection. However, it is assumed that data processing is generally less error-prone.</p>
      </sec>
      <sec id="sec2dot3">
        <title>2.3. Processing Performance</title>
        <p>Processing the data takes some time. The typical tasks take up to a few days to complete. For example, parsing the United States Security Exchange Commission data for the 4th quarter of 2025 (financial filings downloaded from EDGAR on the SEC website; about 3.5 gigabytes of extracted data; please note, this does not include the extraction of the data; only processing the extracted data, which was narrowed down to business, risk factors, legal proceedings, mdna, market risk, and controls procedures in 10-K and 10-Q documents; 5785 files) took a few days on the iFastNet shared hosting platform (2.91 GB RAM, with 3 cores) to build a two-step down table system (2-word table; 3,646,416 rows with 246,093,330 occurrences, and 3-word table; 13,932,144 rows with 358,450,849 occurrences). Parsing the bills for a session of the United States Congress takes under a day (downloaded from the GovInfo website; <ext-link ext-link-type="uri" xlink:href="https://www.govinfo.gov/bulkdata/BILLS/">https://www.govinfo.gov/bulkdata/BILLS/</ext-link>; 330 MB) for a two-step down table system (117th, Session 1, 12,648 files: 2-word table; 983,413 rows with 24,939,993 occurrences, and 3-word; 2,867,280 rows with 21,799,646 occurrences). Note that files were split into sections based upon a period or a newline, cleaned of punctuation, leaving letters, numbers, and hyphens, then split by spaces. No other functions were performed. Adding in parts of speech data to a congressional session bill dataset via the dictionary method takes a few days, although some improvements are needed to the dictionary that was used (Moby Part of Speech List by Grady Ward). Additional abstractions can take a longer amount of time depending upon the complexity of the processing.</p>
      </sec>
      <sec id="sec2dot4">
        <title>2.4. Usefulness</title>
        <p>Even the simplest abstractions produce a significant amount of value and provide a foundation upon which more advanced functionality can be built. At the lowest level, the data provides accurate summarization in navigating datasets, allowing one to perform a full assessment of the dataset as one progresses through it. The conduciveness measures help to understand word and phrase density. Also, anomalies can be identified and readily assessed. Repeating data can also be found with relative ease. A simple webpage was used to load full results and typically returned in under a minute.</p>
        <p>The comparison functionality is very useful. For homogeneous datasets, trends between different historical datasets may be identified. As shown above, positionality may be distinguished between two or more datasets. Conduciveness was included in the testing and gives a sense of the relative density of occurrences, which helps. Parts of speech can be included to ensure better filtering. For example, prepositions and conjunctions may be eliminated from the comparison above. Associating the parts of speech also allows entity comparison. For example, here is an example of verb sets based on a query for president, as partially shown in <bold>Table 5</bold>:</p>
        <p><bold>Table 5</bold><bold>.</bold> Example query results.</p>
        <table-wrap id="tbl5">
          <label>Table 5</label>
          <table>
            <tbody>
              <tr>
                <td>Query</td>
                <td>First Following Word</td>
                <td>Second Following Word</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>submit</td>
              </tr>
              <tr>
                <td>President</td>
                <td>or</td>
                <td>tempore</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>impose</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>provide</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>appoint</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>establish</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>not</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>exercise</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>transmit</td>
              </tr>
              <tr>
                <td>President</td>
                <td>shall</td>
                <td>designate</td>
              </tr>
            </tbody>
          </table>
        </table-wrap>
        <p>As this example demonstrates, it helps to assess the actions that the president shall take. Based on this, word substitutability may be established. For example, submit is similar to impose, provide, appoint, et cetera.</p>
        <p>This provides a basis that may be taken quite a bit further.</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Discussion</title>
      <p>The method outlined here presents a way to analyze and navigate through dataset aggregates and serves as the foundation for more complex tasks. This provides a low barrier to entry into data analysis of larger datasets. Momentum can be captured and extended into more advanced functionality. A database perspective has been presented here, but other abstractions, such as web-based, can be used for further enhancements. Return sets are compressed in a way that allows efficient navigation, even in a server/client model.</p>
      <p>While others utilize data synthesis techniques which require new hardware, this method accommodates existing hardware and offerings. An approach such as this may be more appropriate for many scenarios. Rather than having tools serve as an author based upon the work of others, this method rather allows work to be explored, from which insights may be gained and used in true authoring. It seems that vector searches which pull along only one line of logic, this approach enables the user to gain a more thorough understanding of data, empowering them both in their development and creation.</p>
    </sec>
    <sec id="sec4">
      <title>4. Conclusion</title>
      <p>Here it has been demonstrated how almost anyone may begin participating in semantic aggregation analysis and begin building a foundation for more complex models via a cost-effective method. Utilizing the full-fidelity approach enables more value to be extracted from datasets, lowers the amount of data needed to make a dataset analytically significant, requires less resources to perform the analysis, and provides the foundation for exciting innovation to occur. The democratization of such a practice could yield significant benefits in the field. The more people partaking in this field of study, the more advancements may be produced.</p>
      <sec id="sec4dot1">
        <title>Future Work</title>
        <p>There is much opportunity for innovation, including additional algorithms for data abstractions, domain-specific models, taxonomical standardization, data induction methods, versioned abstraction, integration, security, visualizations, graphing techniques, et cetera. While all of these areas of potential innovation are quite exciting, visualizations may become especially important in multidimensional analytical comparative progressions of datasets; for example, comparing a timelapse of two different datasets for specific models (e.g., weighted analysis of “president” [verb] [word] over the timespan of 1990-2020 for both the United States Security Exchange Commission and the United States Congress to understand how the tasks of presidents compared with each other, and to find likely candidates for future elections). With innovations in graphical rendering and interaction, analysis could become quite rich.</p>
        <p>It was found that the technologies and basis for performing this very basic work were a start, but subpar. Some of the simplest things, like basic words such as the, or, and, et cetera, were mislabeled. Accurate dictionary databases are difficult to come by in the public domain. Academic projects are going up and are being taken down quite frequently. Some of the intellectual property is proprietary. There are also a lot of opportunities for improvement even in the foundational aspects of this type of data analysis.</p>
        <p>It was found that the technologies and basis for performing this very basic work were a start, but subpar. Some of the simplest things, like basic words such as the, or, and, et cetera, were mislabeled. Accurate dictionary databases are difficult to come by in the public domain. Academic projects are going up and are being taken down quite frequently. Some of the intellectual property is proprietary. There are also a lot of opportunities for improvement even in the foundational aspects of this type of data analysis.</p>
      </sec>
    </sec>
    <sec id="sec5">
      <title>Acknowledgement</title>
      <p>Sincere thanks to the members of JDAIP for their professional performance, especially to editorial assistant<italic>Delia Zhu</italic> for achieving collaborative and facilitative excellence. </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Copertari, L.F. (2019) On Natural and Artificial Intelligence. <italic>OALib</italic>, 6, 1-9. https://doi.org/10.4236/oalib.1105221 <pub-id pub-id-type="doi">10.4236/oalib.1105221</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4236/oalib.1105221">https://doi.org/10.4236/oalib.1105221</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Copertari, L.F.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>On Natural and Artificial Intelligence</article-title>
            <source>OALib</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.4236/oalib.1105221</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Duan, J., Zhang, S., Wang, Z., Jiang, L., Qu, W., <italic>et al</italic>. (2024) Efficient Training of Large Language Models on Distributed Infra-Structures: A Survey. arXiv:2407.20018.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Duan, J.</string-name>
              <string-name>Zhang, S.</string-name>
              <string-name>Wang, Z.</string-name>
              <string-name>Jiang, L.</string-name>
              <string-name>Qu, W.</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Efficient Training of Large Language Models on Distributed Infra-Structures: A Survey</article-title>
            <fpage>2407</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Jegham, N., Abdelatti, M., Koh, C.Y., Elmoubarki, L. and Hendawi, A. (2025) How Hungry Is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference. arXiv:2505.09598.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Jegham, N.</string-name>
              <string-name>Abdelatti, M.</string-name>
              <string-name>Koh, C.Y.</string-name>
              <string-name>Elmoubarki, L.</string-name>
              <string-name>Hendawi, A.</string-name>
              <string-name>Energy, W</string-name>
            </person-group>
            <year>2025</year>
            <article-title>How Hungry Is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference</article-title>
            <fpage>2505</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Liu, B., Weng, Z. and Ayinhu (2026) The Advantages and Constraints of LLM-Augmented Knowledge Graphs in University Innovation and Entrepreneurship Education. <italic>Advances</italic><italic>in</italic><italic>Applied</italic><italic>Sociology</italic>, 16, 38-47. https://doi.org/10.4236/aasoci.2026.161003 <pub-id pub-id-type="doi">10.4236/aasoci.2026.161003</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4236/aasoci.2026.161003">https://doi.org/10.4236/aasoci.2026.161003</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Liu, B.</string-name>
              <string-name>Weng, Z.</string-name>
            </person-group>
            <year>2026</year>
            <article-title>The Advantages and Constraints of LLM-Augmented Knowledge Graphs in University Innovation and Entrepreneurship Education</article-title>
            <source>Advances in Applied Sociology</source>
            <volume>16</volume>
            <pub-id pub-id-type="doi">10.4236/aasoci.2026.161003</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wang, X., Tan, Y., Yang, T., Yuan, M., Wang, S., Chen, M., <italic>et al</italic>. (2024) Efficient Large Language Model Application Development: A Case Study of Knowledge Base, API, and Deep Web Search Integration. <italic>Journal</italic><italic>of</italic><italic>Computer</italic><italic>and</italic><italic>Communications</italic>, 12, 171-200. https://doi.org/10.4236/jcc.2024.1212011 <pub-id pub-id-type="doi">10.4236/jcc.2024.1212011</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4236/jcc.2024.1212011">https://doi.org/10.4236/jcc.2024.1212011</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wang, X.</string-name>
              <string-name>Tan, Y.</string-name>
              <string-name>Yang, T.</string-name>
              <string-name>Yuan, M.</string-name>
              <string-name>Wang, S.</string-name>
              <string-name>Chen, M.</string-name>
              <string-name>Base, A</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Efficient Large Language Model Application Development: A Case Study of Knowledge Base, API, and Deep Web Search Integration</article-title>
            <source>Journal of Computer and Communications</source>
            <volume>12</volume>
            <pub-id pub-id-type="doi">10.4236/jcc.2024.1212011</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>