<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ojapps
   </journal-id>
   <journal-title-group>
    <journal-title>
     Open Journal of Applied Sciences
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2165-3917
   </issn>
   <issn publication-format="print">
    2165-3925
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ojapps.2025.1510202
   </article-id>
   <article-id pub-id-type="publisher-id">
    ojapps-146377
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Biomedical 
     </subject>
     <subject>
       Life Sciences, Chemistry 
     </subject>
     <subject>
       Materials Science, Computer Science 
     </subject>
     <subject>
       Communications, Engineering, Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Optimal Predictive Modeling of Nonlinear Transformations: Innovative Applied Mathematics in an Artificial Intelligence System
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Jean Pacifique
      </surname>
      <given-names>
       Nkurunziza
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Fulgence
      </surname>
      <given-names>
       Nahayo
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Ignat
      </surname>
      <given-names>
       Anca
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff3"> 
      <sup>3</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Thierry
      </surname>
      <given-names>
       Nsabimana
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aDoctoral School, Burundi University, Bujumbura, Burundi
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aLURMISTA, ISTA, University of Burundi, Bujumbura, Burundi
    </addr-line> 
   </aff> 
   <aff id="aff3">
    <addr-line>
     aFaculty of Computer Science, University of Alexandru Ion Cuza, Iasi, Rumania
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     30
    </day> 
    <month>
     09
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    15
   </volume> 
   <issue>
    10
   </issue>
   <fpage>
    3073
   </fpage>
   <lpage>
    3093
   </lpage>
   <history>
    <date date-type="received">
     <day>
      5,
     </day>
     <month>
      September
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      12,
     </day>
     <month>
      September
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      12,
     </day>
     <month>
      October
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    In this paper, an Optimal Predictive Modeling of Nonlinear Transformations “OPMNT” method has been developed while using Orthogonal Nonnegative Matrix Factorization “ONMF” with the Kullback-Leibler “KL” divergence. Innovative algorithm includes a Term Frequency-Inverse Document Frequency score smoothing technique and incorporates parameters optimized through cross-validation using the Nelder-Mead method. This technique has numerous applications in Artificial Intelligence Systems (AIS), including extracting nonlinear features for document classification. To make our technique powerful, we apply nonlinear transformations such as logarithms, square roots, and hyperbolic tangents. Our results demonstrate that the OPMNT-KL-ONMF method yields significantly higher accuracy on average compared to untransformed datasets, underscoring the critical importance of selecting appropriate transformation functions to enhance the classification capabilities of the KL-ONMF model. Future work will involve integrating these transformation functions into a neural network framework to explore new techniques for improving performance.
   </abstract>
   <kwd-group> 
    <kwd>
     Optimal Predictive Modeling
    </kwd> 
    <kwd>
      Nonlinear Transformation
    </kwd> 
    <kwd>
      Artificial Intelligence Systems
    </kwd> 
    <kwd>
      Parameter Optimization
    </kwd> 
    <kwd>
      KL-ONMF Model
    </kwd> 
    <kwd>
      Clustering
    </kwd> 
    <kwd>
      Cross-Validation
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>The non-negative matrix factorisation (NMF) method, which was first used in the paper <xref ref-type="bibr" rid="scirp.146377-1">
     [1]
    </xref>, is a technique that allows a non-negative matrix 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math> to be decomposed into two non-negative factors given by 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math>, where</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        X 
      </mi> 
      <mo>
        ≈ 
      </mo> 
      <mi>
        W 
      </mi> 
      <mo>
        × 
      </mo> 
      <mi>
        H 
      </mi> 
      <mo>
        , 
      </mo> 
     </mrow> 
    </math></p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> is an 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        m 
      </mi> 
      <mo>
        × 
      </mo> 
      <mi>
        r 
      </mi> 
     </mrow> 
    </math> matrix with elements in 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msup> 
       <mi>
         ℝ 
       </mi> 
       <mo>
         + 
       </mo> 
      </msup> 
     </mrow> 
    </math>, and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math> is an 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        r 
      </mi> 
      <mo>
        × 
      </mo> 
      <mi>
        n 
      </mi> 
     </mrow> 
    </math> matrix with nonnegative elements. Here, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       m 
     </mi> 
    </math> represents the number of features, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       n 
     </mi> 
    </math> represents the number of observations or samples, while 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       r 
     </mi> 
    </math> denotes the rank or dimensionality of the feature subspace of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math>.</p>
   <p>Let 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mrow> 
       <mo>
         { 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           x 
         </mi> 
         <mn>
           1 
         </mn> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           x 
         </mi> 
         <mn>
           2 
         </mn> 
        </msub> 
        <mo>
          , 
        </mo> 
        <mo>
          ⋯ 
        </mo> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           x 
         </mi> 
         <mi>
           n 
         </mi> 
        </msub> 
       </mrow> 
       <mo>
         } 
       </mo> 
      </mrow> 
      <mo>
        ∈ 
      </mo> 
      <msup> 
       <mi>
         ℝ 
       </mi> 
       <mi>
         m 
       </mi> 
      </msup> 
     </mrow> 
    </math> be a set of data. For any column vector 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         x 
       </mi> 
       <mi>
         j 
       </mi> 
      </msub> 
     </mrow> 
    </math> in 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math>, we can express it as</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         x 
       </mi> 
       <mi>
         j 
       </mi> 
      </msub> 
      <mo>
        ≈ 
      </mo> 
      <munderover> 
       <mstyle mathsize="140%" displaystyle="true"> 
        <mo>
          ∑ 
        </mo> 
       </mstyle> 
       <mrow> 
        <mi>
          k 
        </mi> 
        <mo>
          = 
        </mo> 
        <mn>
          1 
        </mn> 
       </mrow> 
       <mi>
         r 
       </mi> 
      </munderover> 
      <mtext>
          
      </mtext> 
      <msub> 
       <mi>
         w 
       </mi> 
       <mi>
         k 
       </mi> 
      </msub> 
      <msub> 
       <mi>
         h 
       </mi> 
       <mrow> 
        <mi>
          k 
        </mi> 
        <mi>
          j 
        </mi> 
       </mrow> 
      </msub> 
      <mo>
        , 
      </mo> 
     </mrow> 
    </math></p>
   <p>indicating that 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         x 
       </mi> 
       <mi>
         j 
       </mi> 
      </msub> 
     </mrow> 
    </math> can be approximated by a linear combination of the basis vectors (i.e., all column vectors of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math>), with the components of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math> acting as the weight coefficients. Thus, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math> is also known as the coefficient matrix, allowing the feature vector to be computed as a linear combination of the basis vectors.</p>
   <p>Originally proposed for feature extraction, the NMF nonnegative nature makes it suitable for various non-negative processing tasks, underpinned by strong theoretical foundations and interpretability <xref ref-type="bibr" rid="scirp.146377-2">
     [2]
    </xref>. In pattern recognition, dimensionality reduction is crucial for efficiently managing high-dimensional data and improving recognition accuracy. While the K-means algorithm is frequently used for clustering <xref ref-type="bibr" rid="scirp.146377-3">
     [3]
    </xref>, its performance often degrades in high-dimensional spaces, underscoring the necessity for effective dimensionality reduction techniques.</p>
   <p>In addition to pattern recognition, NMF has applications across various fields, including dimension reduction <xref ref-type="bibr" rid="scirp.146377-4">
     [4]
    </xref>, image processing <xref ref-type="bibr" rid="scirp.146377-5">
     [5]
    </xref>, speech processing <xref ref-type="bibr" rid="scirp.146377-6">
     [6]
    </xref> <xref ref-type="bibr" rid="scirp.146377-7">
     [7]
    </xref>, spectral analysis <xref ref-type="bibr" rid="scirp.146377-8">
     [8]
    </xref> <xref ref-type="bibr" rid="scirp.146377-9">
     [9]
    </xref>, DNA expression analysis <xref ref-type="bibr" rid="scirp.146377-10">
     [10]
    </xref> <xref ref-type="bibr" rid="scirp.146377-11">
     [11]
    </xref>, microRNA-disease analysis <xref ref-type="bibr" rid="scirp.146377-12">
     [12]
    </xref> <xref ref-type="bibr" rid="scirp.146377-13">
     [13]
    </xref>, social network analysis <xref ref-type="bibr" rid="scirp.146377-14">
     [14]
    </xref>, text analysis <xref ref-type="bibr" rid="scirp.146377-15">
     [15]
    </xref>, hyperspectral signal unmixing <xref ref-type="bibr" rid="scirp.146377-16">
     [16]
    </xref>, and blind spectral unmixing <xref ref-type="bibr" rid="scirp.146377-17">
     [17]
    </xref> <xref ref-type="bibr" rid="scirp.146377-18">
     [18]
    </xref>. As data size increases, so does the demand for processing and storage. This has led researchers to develop new algorithms for linear dimensionality reduction (LDR), a data analysis method that reduces the dimensionality of information with minimal loss of information before analysis. The transformation of data is an effective means to improve the linear representation. The idea is to transform the data via a function 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       f 
     </mi> 
    </math> to be learned, and use 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        f 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         X 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mrow> 
       <mo>
         [ 
       </mo> 
       <mrow> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <mo>
           ( 
         </mo> 
         <mrow> 
          <msub> 
           <mi>
             θ 
           </mi> 
           <mn>
             1 
           </mn> 
          </msub> 
         </mrow> 
         <mo>
           ) 
         </mo> 
        </mrow> 
        <mo>
          , 
        </mo> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <mo>
           ( 
         </mo> 
         <mrow> 
          <msub> 
           <mi>
             θ 
           </mi> 
           <mn>
             2 
           </mn> 
          </msub> 
         </mrow> 
         <mo>
           ) 
         </mo> 
        </mrow> 
        <mo>
          , 
        </mo> 
        <mo>
          ⋯ 
        </mo> 
        <mo>
          , 
        </mo> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <mo>
           ( 
         </mo> 
         <mrow> 
          <msub> 
           <mi>
             θ 
           </mi> 
           <mi>
             n 
           </mi> 
          </msub> 
         </mrow> 
         <mo>
           ) 
         </mo> 
        </mrow> 
       </mrow> 
       <mo>
         ] 
       </mo> 
      </mrow> 
     </mrow> 
    </math> and apply the factorization to 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        f 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         X 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> instead of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math>, i.e. 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        f 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         X 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        ≈ 
      </mo> 
      <mi>
        W 
      </mi> 
      <mo>
        × 
      </mo> 
      <mi>
        H 
      </mi> 
     </mrow> 
    </math>.</p>
   <p>We assume that each data point is associated with a label 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         y 
       </mi> 
       <mn>
         1 
       </mn> 
      </msub> 
      <mo>
        , 
      </mo> 
      <mo>
        ⋯ 
      </mo> 
      <mo>
        , 
      </mo> 
      <msub> 
       <mi>
         y 
       </mi> 
       <mi>
         n 
       </mi> 
      </msub> 
     </mrow> 
    </math>. In the most basic case, they have labels 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         y 
       </mi> 
       <mi>
         i 
       </mi> 
      </msub> 
     </mrow> 
    </math> that are binary, i.e., they can take two values representing two classes. For multi-class cases, the labels can be represented using one-hot encoding <xref ref-type="bibr" rid="scirp.146377-19">
     [19]
    </xref>. A prominent example of a non-linear transformation applied to count data in document classification is the term frequency-inverse document frequency (TF-IDF) transformation <xref ref-type="bibr" rid="scirp.146377-20">
     [20]
    </xref> <xref ref-type="bibr" rid="scirp.146377-21">
     [21]
    </xref>. In this context, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        X 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> denotes the frequency of word 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       i 
     </mi> 
    </math> appearing in document 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       j 
     </mi> 
    </math>, and the TF-IDF transformation maps 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math> to 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        f 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         X 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> where</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        X 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mi>
        f 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          X 
        </mi> 
        <mrow> 
         <mo>
           ( 
         </mo> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mo>
            , 
          </mo> 
          <mi>
            j 
          </mi> 
         </mrow> 
         <mo>
           ) 
         </mo> 
        </mrow> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        . 
      </mo> 
     </mrow> 
    </math></p>
   <p>In <xref ref-type="bibr" rid="scirp.146377-22">
     [22]
    </xref> <xref ref-type="bibr" rid="scirp.146377-23">
     [23]
    </xref>, document classification is achieved through NMF and KL-NMF. Additionally, <xref ref-type="bibr" rid="scirp.146377-24">
     [24]
    </xref> demonstrates the application of non-linear transformations to enhance the effectiveness of these classification methods. This motivates our exploration of improved transformations, which we aim to apply to various document datasets. This motivates our exploration of some nonlinear transformations, which we intend to apply to various document datasets to enhance the results obtained in <xref ref-type="bibr" rid="scirp.146377-25">
     [25]
    </xref>.</p>
   <p>Another effective method is the probabilistic IDF <xref ref-type="bibr" rid="scirp.146377-26">
     [26]
    </xref>, which modifies the basic IDF calculation to incorporate probabilities, ensuring that terms that do not appear in a document are still assigned meaningful weights. Jelinek-Mercer smoothing <xref ref-type="bibr" rid="scirp.146377-27">
     [27]
    </xref> <xref ref-type="bibr" rid="scirp.146377-28">
     [28]
    </xref>, initially developed for language modeling, can also be adapted for TF-IDF by interpolating maximum likelihood estimates with background distributions.</p>
   <p>Moreover, Katz <xref ref-type="bibr" rid="scirp.146377-29">
     [29]
    </xref> and Witten-Bell <xref ref-type="bibr" rid="scirp.146377-30">
     [30]
    </xref> smoothing techniques provide further enhancements by allowing for interpolation between observed and unobserved frequencies, improving robustness in document classification tasks at language analysis. It is important to emphasize the use of these smoothing techniques to improve the performance of TF-IDF so that the weights assigned to terms indicate the general importance of the terms in the overall corpus.</p>
   <p>Outline and Contribution of the Paper</p>
   <p>This paper proposes a new method for document classification. As detailed in Section 2, KL-ONMF serves as the reference for our method. We will smooth the TF-IDF and incorporate parameters (see Section 3), which we will optimize using the fminsearch function in MATLAB with a machine learning approach. We apply non-linear transformations in order to capture non-linear characteristics better. We will show that this new approach is much better than KL-ONMF on the original data for document classification (see Section 4).</p>
  </sec><sec id="s2">
   <title>2. Alternating Optimization for ONMF with the KL Divergence</title>
   <p>In this section, we focus on Alternating Optimization for Orthogonal Nonnegative Matrix Factorization (ONMF) with the Kullback-Leibler (KL) Divergence. The goal is to minimize the KL Divergence between the given matrix 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math> and the product of two matrices 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math>, while enforcing specific constraints on 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math>.</p>
   <p>The optimization problem can be formulated as follows:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <munder> 
       <mrow> 
        <mtext>
          min 
        </mtext> 
       </mrow> 
       <mrow> 
        <mi>
          W 
        </mi> 
        <mo>
          ∈ 
        </mo> 
        <msup> 
         <mi>
           ℝ 
         </mi> 
         <mrow> 
          <mi>
            m 
          </mi> 
          <mo>
            × 
          </mo> 
          <mi>
            r 
          </mi> 
         </mrow> 
        </msup> 
        <mo>
          , 
        </mo> 
        <mi>
          H 
        </mi> 
        <mo>
          ∈ 
        </mo> 
        <msup> 
         <mi>
           ℝ 
         </mi> 
         <mrow> 
          <mi>
            r 
          </mi> 
          <mo>
            × 
          </mo> 
          <mi>
            n 
          </mi> 
         </mrow> 
        </msup> 
       </mrow> 
      </munder> 
      <msub> 
       <mi>
         D 
       </mi> 
       <mrow> 
        <mtext>
          KL 
        </mtext> 
       </mrow> 
      </msub> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          X 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          W 
        </mi> 
        <mi>
          H 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mtext>
          
      </mtext> 
      <mtext>
          
      </mtext> 
      <mtext>
        subject 
      </mtext> 
      <mtext>
          
      </mtext> 
      <mtext>
        to 
      </mtext> 
      <mtext>
          
      </mtext> 
      <mtext>
          
      </mtext> 
      <mi>
        H 
      </mi> 
      <mo>
        ≥ 
      </mo> 
      <mn>
        0 
      </mn> 
      <mtext>
          
      </mtext> 
      <mtext>
          
      </mtext> 
      <mtext>
        and 
      </mtext> 
      <mtext>
          
      </mtext> 
      <mtext>
          
      </mtext> 
      <mi>
        H 
      </mi> 
      <msup> 
       <mi>
         H 
       </mi> 
       <mo>
         ⊤ 
       </mo> 
      </msup> 
      <mo>
        = 
      </mo> 
      <msub> 
       <mi>
         I 
       </mi> 
       <mi>
         r 
       </mi> 
      </msub> 
      <mo>
        . 
      </mo> 
     </mrow> 
    </math></p>
   <p>For ONMF utilizing the KL Divergence, it is crucial that 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math> is component-wise nonnegative, reflecting the inherent properties of KL Divergence, which is only defined for non-negative matrices. This non-negativity is essential, as it ensures that the divergence measurement remains valid and interpretable in the context of probability distributions.</p>
   <p>In contrast, ONMF using the Frobenius norm does not impose such restrictions, allowing for broader applications but potentially leading to less interpretable results in probabilistic terms. The choice of KL Divergence in this context emphasizes our focus on modeling non-negative data, making it particularly suitable for applications like document classification or image processing, where the underlying data naturally adheres to non-negativity constraints.</p>
   <p>Moreover, the constraint 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        H 
      </mi> 
      <msup> 
       <mi>
         H 
       </mi> 
       <mo>
         ⊤ 
       </mo> 
      </msup> 
      <mo>
        = 
      </mo> 
      <msub> 
       <mi>
         I 
       </mi> 
       <mi>
         r 
       </mi> 
      </msub> 
     </mrow> 
    </math> ensures that the matrix 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math> retains orthogonality, which is vital for preserving the interpretability of the factors derived from the decomposition. This orthogonality condition aids in distinctly separating the components, allowing for clearer insights into the latent structures represented by the factors in 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math>.</p>
   <p>This formulation highlights the methodological rigor behind ONMF with KL Divergence, ensuring that the results are both mathematically sound and practically applicable to real-world datasets. The update rules for 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math> in ONMF with the KL Divergence highlight the different optimization strategies compared to Frobenius ONMF, making KL-ONMF less sensitive to outliers and more focused on data points with smaller norms.</p>
   <p>Algorithm 1 summarizes the alternating optimization scheme, known as KL-ONMF.</p>
   <fig id="fig1" position="float">
    <label>Figure 1</label>
    <caption>
     <title>It’s a simple yet effective and highly scalable alternating optimization algorithm. It’s an alternating optimization algorithm, which is simple but efficient and highly scalable, running in 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mi>
         
   O
  
        </mi>
  
        <mrow>
   
         <mo>
          
    (
   
         </mo> 
   
         <mrow> 
    
          <mtext>
           
     nnz
    
          </mtext>
    
          <mrow>
     
           <mo>
             ( 
           </mo> 
     
           <mi>
             X 
           </mi> 
     
           <mo>
             ) 
           </mo>
    
          </mrow>
    
          <mi>
           
     r
    
          </mi>
   
         </mrow> 
   
         <mo>
          
    )
   
         </mo>
  
        </mrow>
 
       </mrow>

      </math> operations where 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mtext>
         
   nnz
  
        </mtext>
  
        <mrow>
   
         <mo>
          
    (
   
         </mo> 
   
         <mi>
          
    X
   
         </mi> 
   
         <mo>
          
    )
   
         </mo>
  
        </mrow>
 
       </mrow>

      </math> is the number of non-zero entries in the data matrix, and 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  r
 
       </mi>

      </math> is the number of clusters. Applied on documents and hyperspectral images, KL-ONMF demonstrates its performance over ONMF using the Frobenius norm by providing better clustering results, and the convergence time is always smaller on average.3. Optimal Predictive Modeling of Nonlinear TransformationsIn this section, we develop smoothing techniques to enhance the calculations of Term Frequency-Inverse Document Frequency (TF-IDF). These techniques address limitations associated with raw term frequency counts, ensuring a more balanced representation of term importance.3.1. TF-IDF ModelingIn <xref ref-type="bibr" rid="scirp.146377-31">
       [31]
      </xref>, the authors define TF-IDF as a combination of two statistics: term frequency (TF) and inverse document frequency (IDF). The term frequency 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mi>
         
   t
  
        </mi>
  
        <mi>
         
   f
  
        </mi>
  
        <mrow>
   
         <mo>
          
    (
   
         </mo> 
   
         <mrow> 
    
          <mi>
           
     i
    
          </mi>
    
          <mo>
           
     ,
    
          </mo>
    
          <mi>
           
     j
    
          </mi>
   
         </mrow> 
   
         <mo>
          
    )
   
         </mo>
  
        </mrow>
 
       </mrow>

      </math> quantifies how often term 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  i
 
       </mi>

      </math> appears in document 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  j
 
       </mi>

      </math> <xref ref-type="bibr" rid="scirp.146377-32">
       [32]
      </xref>, and is expressed as:
      <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mi>
         
   t
  
        </mi>
  
        <mi>
         
   f
  
        </mi>
  
        <mrow>
   
         <mo>
          
    (
   
         </mo> 
   
         <mrow> 
    
          <mi>
           
     i
    
          </mi>
    
          <mo>
           
     ,
    
          </mo>
    
          <mi>
           
     j
    
          </mi>
   
         </mrow> 
   
         <mo>
          
    )
   
         </mo>
  
        </mrow>
  
        <mo>
         
   =
  
        </mo>
  
        <mfrac> 
   
         <mrow> 
    
          <msub> 
     
           <mi>
             f 
           </mi> 
     
           <mrow> 
            <mi>
              i 
            </mi> 
            <mo>
              , 
            </mo> 
            <mi>
              j 
            </mi> 
           </mrow> 
    
          </msub> 
   
         </mrow> 
   
         <mrow> 
    
          <mstyle displaystyle="true"> 
     
           <munder> 
            <mo>
              ∑ 
            </mo> 
            <mrow> 
             <msup> 
              <mi>
                i 
              </mi> 
              <mo>
                ′ 
              </mo> 
             </msup> 
             <mo>
               ∈ 
             </mo> 
             <mi>
               j 
             </mi> 
            </mrow> 
           </munder> 
     
           <mrow> 
            <msub> 
             <mi>
               f 
             </mi> 
             <mrow> 
              <msup> 
               <mi>
                 i 
               </mi> 
               <mo>
                 ′ 
               </mo> 
              </msup> 
              <mo>
                , 
              </mo> 
              <mi>
                j 
              </mi> 
             </mrow> 
            </msub> 
           </mrow> 
    
          </mstyle>
   
         </mrow> 
  
        </mfrac> 
  
        <mo>
         
   .
  
        </mo>
 
       </mrow>

      </math>(1)where 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <msub> 
   
         <mi>
          
    f
   
         </mi> 
   
         <mrow> 
    
          <mi>
           
     i
    
          </mi>
    
          <mo>
           
     ,
    
          </mo>
    
          <mi>
           
     j
    
          </mi>
   
         </mrow> 
  
        </msub> 
 
       </mrow>

      </math> is the raw count of the term in the document, and 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mstyle displaystyle="true"> 
   
         <munder> 
    
          <mo>
           
     ∑
    
          </mo> 
    
          <mrow> 
     
           <msup> 
            <mi>
              i 
            </mi> 
            <mo>
              ′ 
            </mo> 
           </msup> 
     
           <mo>
             ∈ 
           </mo>
     
           <mi>
             j 
           </mi>
    
          </mrow> 
   
         </munder> 
   
         <mrow> 
    
          <msub> 
     
           <mi>
             f 
           </mi> 
     
           <mrow> 
            <msup> 
             <mi>
               i 
             </mi> 
             <mo>
               ′ 
             </mo> 
            </msup> 
            <mo>
              , 
            </mo> 
            <mi>
              j 
            </mi> 
           </mrow> 
    
          </msub> 
   
         </mrow> 
  
        </mstyle>
 
       </mrow>

      </math> is the total number of terms in document 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  j
 
       </mi>

      </math>. Inverse document frequency measures the amount of information provided by a word, indicating whether it is present in all documents, and is defined as the inverse fraction, on a logarithmic scale, of documents that contain the word. It is given by:
      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mi>
         
   i
  
        </mi>
  
        <mi>
         
   d
  
        </mi>
  
        <mi>
         
   f
  
        </mi>
  
        <mrow>
   
         <mo>
          
    (
   
         </mo> 
   
         <mi>
          
    i
   
         </mi> 
   
         <mo>
          
    )
   
         </mo>
  
        </mrow>
  
        <mo>
         
   =
  
        </mo>
  
        <mtext>
         
   log
  
        </mtext>
  
        <mrow>
   
         <mo>
          
    (
   
         </mo> 
   
         <mrow> 
    
          <mfrac> 
     
           <mi>
             N 
           </mi> 
     
           <mrow> 
            <mrow> 
             <mo>
               | 
             </mo> 
             <mrow> 
              <mrow> 
               <mo>
                 { 
               </mo> 
               <mrow> 
                <mi>
                  j 
                </mi> 
                <mo>
                  ∈ 
                </mo> 
                <mi>
                  J 
                </mi> 
                <mo>
                  : 
                </mo> 
                <mi>
                  i 
                </mi> 
                <mo>
                  ∈ 
                </mo> 
                <mi>
                  j 
                </mi> 
               </mrow> 
               <mo>
                 } 
               </mo> 
              </mrow> 
             </mrow> 
             <mo>
               | 
             </mo> 
            </mrow> 
           </mrow> 
    
          </mfrac> 
   
         </mrow> 
   
         <mo>
          
    )
   
         </mo>
  
        </mrow>
  
        <mo>
         
   .
  
        </mo>
 
       </mrow>

      </math>(2)where 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  N
 
       </mi>

      </math> is the total number of documents in the corpus, and 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <mrow>
   
         <mo>
          
    |
   
         </mo> 
   
         <mi>
          
    j
   
         </mi> 
   
         <mo>
          
    |
   
         </mo>
  
        </mrow>
 
       </mrow>

      </math> is the total number of terms in document 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  j
 
       </mi>

      </math>. This approach effectively balances the frequency and rarity of terms, making it a powerful tool for information retrieval and text analysis. In the next section, we will introduce smoothing techniques to enhance the TF-IDF representation, aiming to improve the robustness and performance of our classification method.3.2. Smoothing Technique for TF-IDF in Optimal Predictive Modeling</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2313390-rId107.jpeg?20251015031439" />
   </fig>
   <p>To address the limitations of raw term frequency counts and ensure a more balanced representation of term importance, we introduce smoothing techniques in the calculation of both term frequency (TF) and inverse document frequency (IDF). The term frequency is recalculated with a smoothing term to prevent zero counts from disproportionately affecting the results. Similarly, the inverse document frequency is adjusted to ensure that terms common across many documents do not dominate the weighting.</p>
   <p>The term frequency 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mtext>
        TF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> is computed as:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mtext>
        TF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mfrac> 
       <mrow> 
        <msub> 
         <mi>
           f 
         </mi> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mo>
            , 
          </mo> 
          <mi>
            j 
          </mi> 
         </mrow> 
        </msub> 
        <mo>
          + 
        </mo> 
        <mtext>
          smoothing 
        </mtext> 
       </mrow> 
       <mrow> 
        <mstyle displaystyle="true"> 
         <munder> 
          <mo>
            ∑ 
          </mo> 
          <mrow> 
           <msup> 
            <mi>
              i 
            </mi> 
            <mo>
              ′ 
            </mo> 
           </msup> 
           <mo>
             ∈ 
           </mo> 
           <mi>
             j 
           </mi> 
          </mrow> 
         </munder> 
         <mrow> 
          <msub> 
           <mi>
             f 
           </mi> 
           <mrow> 
            <msup> 
             <mi>
               i 
             </mi> 
             <mo>
               ′ 
             </mo> 
            </msup> 
            <mo>
              , 
            </mo> 
            <mi>
              j 
            </mi> 
           </mrow> 
          </msub> 
         </mrow> 
        </mstyle> 
        <mo>
          + 
        </mo> 
        <mtext>
          smoothing 
        </mtext> 
       </mrow> 
      </mfrac> 
      <mo>
        , 
      </mo> 
     </mrow> 
    </math>(3)</p>
   <p>where smooth = 1 for Laplace smoothing or add-1 smoothing <xref ref-type="bibr" rid="scirp.146377-33">
     [33]
    </xref>. The inverse document frequency 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mtext>
        IDF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         i 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> is calculated to measure the importance of the term across the document corpus:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mtext>
        IDF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         i 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mtext>
        log 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mfrac> 
         <mrow> 
          <mi>
            N 
          </mi> 
          <mo>
            + 
          </mo> 
          <mtext>
            smoothing 
          </mtext> 
         </mrow> 
         <mrow> 
          <mtext>
            DF 
          </mtext> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             i 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mo>
            + 
          </mo> 
          <mtext>
            smoothing 
          </mtext> 
         </mrow> 
        </mfrac> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        . 
      </mo> 
     </mrow> 
    </math>(4)</p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mtext>
        DF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         i 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> represents the number of documents containing the term 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       i 
     </mi> 
    </math>. This formulation helps to mitigate the impact of terms that appear in a large number of documents, ensuring that the IDF value remains informative. The TF-IDF weighting of a term is therefore the product of its TF term frequency Equation (3) and its IDF document inverse frequency Equation (4). It is defined by Equation (5):</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        X 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <mtext>
        TF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        × 
      </mo> 
      <mtext>
        IDF 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mi>
         i 
       </mi> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        . 
      </mo> 
     </mrow> 
    </math>(5)</p>
   <p>In the following, we start with a matrix 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math> representing word occurrences, where each element 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        X 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          , 
        </mo> 
        <mi>
          j 
        </mi> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> indicates how many times word 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       i 
     </mi> 
    </math> appears in document 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       j 
     </mi> 
    </math>. The algorithm computes the Term Frequency (TF) for each word in a document, reflecting its relative importance, and the Inverse Document Frequency (IDF) measures how common or rare a word is across all documents. The resulting TF-IDF scores highlight the significance of words in the context of the documents. The detailed algorithm is described below.</p>
   <sec id="s2_1">
    <title>3.3. Parameterization of TF-IDF Score with Custom Adjustments</title>
    <p>To further refine the TF-IDF score, we introduce adjustable parameters 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        a 
      </mi> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        b 
      </mi> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        c 
      </mi> 
     </math>. These parameters enable customization based on the dataset’s characteristics, thereby enhancing the overall performance of clustering algorithms. The TF-IDF score incorporating parameters to be optimized can be expressed as follows:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         X 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           j 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         a 
       </mi> 
       <mo>
         ⋅ 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           X 
         </mi> 
         <msup> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               i 
             </mi> 
             <mo>
               , 
             </mo> 
             <mi>
               j 
             </mi> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mi>
            b 
          </mi> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         c 
       </mi> 
       <mo>
         ⋅ 
       </mo> 
       <mi>
         X 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           j 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math>(6)</p>
    <p>We note that for the classic TF-IDF, we have 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         a 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         b 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         c 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> or 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         a 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         b 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         c 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
      </mrow> 
     </math>. In our new model, each parameter plays a crucial role in fine-tuning the model to capture the underlying patterns in the data: parameter 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        a 
      </mi> 
     </math> scales the contribution of the transformed term frequencies. By adjusting 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        a 
      </mi> 
     </math>, we can amplify the significance of certain terms based on their frequency, ensuring that important terms are more prominently represented in the final feature set. Parameter 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        b 
      </mi> 
     </math> serves as the exponent applied to the term frequency matrix 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        X 
      </mi> 
     </math>. This exponentiation can accentuate the differences among term frequencies, effectively emphasizing more frequent terms while diminishing the influence of less common ones, and parameter 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        c 
      </mi> 
     </math> adjusts the influence of the original TF-IDF scores.</p>
   </sec>
   <sec id="s2_2">
    <title>3.4. Innovative Nonlinear Transformations: Boosting Predictive Performance in AI Applications</title>
    <p>In this section, we explore three nonlinear transformations that enhance the model’s ability to capture complex relationships within data. These transformations are inspired by principles from neural networks and include the logarithm function, the square root function, and the hyperbolic tangent function. Each transformation serves a specific purpose in refining the feature representation of TF-IDF scores.</p>
    <p>Logarithm Function</p>
    <p>We prefer the logarithm function due to its effectiveness in reducing skewness in data distributions. Many real-world datasets exhibit right skewness; applying a logarithmic transformation can render the data more symmetric, thereby improving the performance of clustering algorithms. The logarithmic transformation can be expressed as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mtext>
         log 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           + 
         </mo> 
         <mi>
           x 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math>(7)</p>
    <p>Moreover, logarithmic scaling emphasizes smaller differences in low-value ranges, which is particularly beneficial for TF-IDF scores. This allows for a more nuanced understanding of infrequent terms that may still hold significant semantic weight. Additionally, by compressing the range of values, logarithmic transformations mitigate the influence of outliers, making clustering algorithms less sensitive to extreme values.</p>
    <p>Square Root Function</p>
    <p>The square root function is particularly advantageous for processing non-negative TF-IDF scores, as it emphasizes larger values while preserving the non-negativity of the data. This property is crucial for highlighting important terms that appear more frequently within the dataset, thereby amplifying their significance in the analysis. We use the expression given by:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <msqrt> 
        <mi>
          x 
        </mi> 
       </msqrt> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math>(8)</p>
    <p>Moreover, the square root transformation effectively mitigates the impact of extremely high values, enabling clustering algorithms to concentrate on the overall structure of the data rather than being skewed by a few dominant terms. Since many clustering methods rely on distance measures, such as Euclidean distance, the square root transformation helps to maintain the relative distances between data points. This results in more accurate and meaningful clustering outcomes, enhancing the model’s ability to identify distinct groups within the data.</p>
    <p>Hyperbolic Tangent Function</p>
    <p>The hyperbolic tangent function outputs values in the range of −1 to 1, which is particularly beneficial for clustering algorithms that assume centered data. This symmetry facilitates the model’s ability to learn balanced representations of both positive and negative influences within the dataset.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mtext>
         tanh 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mi>
            x 
          </mi> 
         </msup> 
         <mo>
           − 
         </mo> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <mi>
             x 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
        <mrow> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mi>
            x 
          </mi> 
         </msup> 
         <mo>
           + 
         </mo> 
         <msup> 
          <mtext>
            e 
          </mtext> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <mi>
             x 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
       </mfrac> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math>(9)</p>
    <p>Additionally, the tanh function introduces strong nonlinearity, which aids in capturing complex relationships among data points. This capability is crucial in clustering, where relationships may not be linearly separable. In the context of neural networks, the tanh function can also lead to more efficient gradients during backpropagation, potentially improving convergence rates in learning models.</p>
    <p>In summary, we have chosen these transformations because they introduce nonlinearity into the model, allowing it to better learn from the underlying patterns in the data. The effectiveness of these nonlinear transformations is further enhanced by the previously optimized parameters 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        a 
      </mi> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        b 
      </mi> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        c 
      </mi> 
     </math>. By refining the representation of term importance through parameter tuning, we ensure that these transformations can operate on the optimized TF-IDF scores 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        x 
      </mi> 
     </math>, resulting in a more relevant and expressive feature space. This integrated approach—comprising smoothing, parameter optimization, and nonlinear transformations—aims to enhance clustering accuracy and improve document classification and retrieval outcomes.</p>
   </sec>
   <sec id="s2_3">
    <title>3.5. Algorithm: Optimizing TF-IDF Score with Smoothing Techniques</title>
    <p>Algorithm 2 summarizes the alternating optimization scheme, which we refer to as KL-ONMF. The index 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        i 
      </mi> 
     </math> refers of terms in the TF-IDF vector (ranging from 1 to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        n 
      </mi> 
     </math>), and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        j 
      </mi> 
     </math> is the index of documents (for each term 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        j 
      </mi> 
     </math> varies based on the number of documents).</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Algorithm 2 is designed to enhance the traditional TF-IDF scoring method by incorporating smoothing and nonlinear transformations. It takes as input a vector of TF-IDF scores along with several parameters, including smoothing and scaling coefficient 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  a
 
        </mi>

       </math>, exponent 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  b
 
        </mi>

       </math>, and adjustment factor 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  c
 
        </mi>

       </math>, with the goal to produce an optimized vector of TF-IDF scores and corresponding cluster labels.3.6. Parameter Optimization Using RegularizationTo identify optimal values for parameters 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  b
 
        </mi>

       </math> and 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  c
 
        </mi>

       </math>, we leverage the fminsearch function in MATLAB, which performs unconstrained optimization. During this process, the parameter 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  a
 
        </mi>

       </math> is fixed at a baseline value of 0.0001. Initial values for 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  b
 
        </mi>

       </math> and 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  c
 
        </mi>

       </math> are generated randomly within a defined range of 0 to 2, enabling a thorough exploration of the parameter space and facilitating the discovery of improved configurations.In our optimization function, we incorporate a regularization term to mitigate the risk of overfitting. This term penalizes excessively large values of the parameters 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  b
 
        </mi>

       </math> and 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
         
  c
 
        </mi>

       </math> by introducing a quadratic penalty, represented as 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
         <mi>
          
   λ
  
         </mi>
  
         <mstyle displaystyle="true"> 
   
          <munderover> 
    
           <mo>
            
     ∑
    
           </mo> 
    
           <mrow> 
     
            <mi>
              i 
            </mi>
     
            <mo>
              = 
            </mo>
     
            <mn>
              1 
            </mn>
    
           </mrow> 
    
           <mi>
            
     k
    
           </mi> 
   
          </munderover> 
   
          <mrow> 
    
           <msubsup> 
     
            <mi>
              θ 
            </mi> 
     
            <mi>
              i 
            </mi> 
     
            <mn>
              2 
            </mn> 
    
           </msubsup> 
   
          </mrow> 
  
         </mstyle>
 
        </mrow>

       </math>, where 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
         <mi>
          
   k
  
         </mi>
  
         <mo>
          
   =
  
         </mo>
  
         <mn>
          
   2
  
         </mn>
 
        </mrow>

       </math> denotes the number of parameters. Specifically, this is an L2 regularization <xref ref-type="bibr" rid="scirp.146377-34">
        [34]
       </xref> component that helps to prevent overfitting by penalizing large parameter values, with each 

       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
         <msub> 
   
          <mi>
           
    θ
   
          </mi> 
   
          <mi>
           
    i
   
          </mi> 
  
         </msub> 
 
        </mrow>

       </math> representing a parameter being optimized. By doing so, we encourage the model to prioritize simpler configurations that generalize better to unseen data. The inclusion of this regularization term helps maintain a balance between fitting the training data and preserving the model’s ability to perform well on validation sets, ultimately leading to more robust and reliable results.3.7. Data SeparationIn order to evaluate the robustness of the optimized methodology, it is essential to separate the data into training and test sets. This separation is crucial for assessing the model’s ability to generalize to unseen data and is performed by allocating 80% of the documents to the training set and 20% to the validation set.The 80/20 split is strategically chosen as it strikes a balance between providing ample data for training the model and retaining a substantial portion for evaluation. This division ensures that the model can effectively learn from the training data while maintaining the capability to accurately assess its performance on independent data. Such a balance is crucial for developing a robust model that can generalize well to new, unseen instances.In this context, “training” refers to the phase in which the model actively learns from the training set, adjusting its parameters to minimize prediction errors among terms and their corresponding clusters. Conversely, “validation” is the process of evaluating the model’s performance on unseen data, which is essential for understanding its generalization capabilities. This two-phase approach allows us to ensure that the model is not only fitting the training data well but is also capable of making accurate predictions in real-world scenarios.Furthermore, the use of labels in this unsupervised task is limited strictly to evaluation purposes. During the training phase, labels are not utilized; instead, they serve as a benchmark for measuring the accuracy and effectiveness of the model’s clustering outcomes. This methodology emphasizes our commitment to an unbiased evaluation process, ensuring that the model’s performance is assessed based solely on its ability to identify and group similar documents without any prior knowledge of their labels.To enhance the evaluation process, we also employ cross-validation. Specifically, we set the number of folds for cross-validation to num_folds and the number of repetitions to num_repetitions.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/2313390-rId208.jpeg?20251015031441" />
    </fig>
    <p>Cross-validation is a resampling technique that helps in assessing how the results of a statistical analysis will generalize to an independent dataset. By splitting the training set into several subsets (folds), we can train the model on a portion of the data while validating it on the remaining parts. This method is repeated for a specified number of folds, providing a more comprehensive evaluation of model performance.</p>
    <p>1) Number of Folds: Setting num_folds = 5 means that the training data will be split into five subsets. The model will be trained on one subset and validated on the other, and this process will be repeated for each fold.</p>
    <p>2) Number of Repetitions: With num_repetitions = 5, the entire cross-validation process will be repeated five times, ensuring that the model’s performance is robust and less sensitive to the specific data split.</p>
    <p>By using cross-validation, we can achieve a more reliable assessment of our model’s performance, reducing the risk of overfitting and ensuring that the enhancements made to the TF-IDF calculations yield tangible benefits in practical applications.</p>
   </sec>
  </sec><sec id="s3">
   <title>4. Numerical Experiments</title>
   <p>In this section, we use nonlinear functions for document clustering, and we compare the performance of KL-ONMF (Algorithm developed in <xref ref-type="bibr" rid="scirp.146377-25">
     [25]
    </xref>) applied to the original data with that of KL-ONMF applied to the transformed data using (Algorithm). All experiments were run on a LAPTOP 11<sup>th</sup> Gen Intel<sup>®</sup> Core<sup>TM</sup> i5-1135G7 @ 2.40GHz × 8 16,0 Go RAM.</p>
   <p>Initialization:</p>
   <p>Similar to k-means, ONMF algorithms can be initialized in a variety of ways. We take the approach suggested in <xref ref-type="bibr" rid="scirp.146377-35">
     [35]
    </xref> to simplify the presentation. The successive nonnegative projection algorithm (SNPA) <xref ref-type="bibr" rid="scirp.146377-36">
     [36]
    </xref> is used in this method to initialize 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math>. It finds a subset of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       r 
     </mi> 
    </math> columns from 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       X 
     </mi> 
    </math> that accurately represent the dataset’s evenly distributed data points.</p>
   <p>To refine the parameters during the optimization phase, we utilize the Nelder-Mead optimization algorithm <xref ref-type="bibr" rid="scirp.146377-37">
     [37]
    </xref>. For this algorithm, we initialize the parameters by generating random values. The Nelder-Mead algorithm iteratively refines these parameters based on the objective function (Equation (10)), which in our case is designed to minimize the difference between the transformed data and the original data, while also incorporating a regularization term. By combining the SNPA for the initialization of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> and the Nelder-Mead method for parameter optimization, we aim to achieve robust and efficient convergence in the ONMF framework.</p>
   <p>Parameterization of Functions:</p>
   <p>To minimize the error rate of the objective functions and enhance accuracy, we begin by setting a fixed initial parameter value for 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> at 0.0001. This choice ensures consistency across our optimization process, allowing us to focus on adjusting the other parameters, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math>. The values for 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> are randomly initialized (see Section 3.6), specifically rounded to three decimal places to maintain precision. This randomness introduces variability in the initial conditions, which can help the optimization algorithm explore different configurations effectively. To find the optimal values of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math>, we employ the fminsearch algorithm, which utilizes the Nelder-Mead method. This optimization technique is particularly beneficial for multidimensional problems where derivatives may not be easily computable. It iteratively adjusts the parameters based on the evaluation of the objective function, seeking to minimize the error rate. By systematically refining the values of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math>, the algorithm converges towards a local parameterization that improves the model’s accuracy in its predictions. Overall, this approach ensures a robust framework for optimizing the model’s performance based on the given dataset.</p>
   <p>Objective Function:</p>
   <p>The objective function aims to minimize the sum of the squared deviations between the calculated and observed outflows. This approach ensures that the model’s predictions closely align with the actual data. Let 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       Θ 
     </mi> 
    </math> be the set of parameters defined as:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        Θ 
      </mi> 
      <mo>
        = 
      </mo> 
      <mrow> 
       <mo>
         { 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           1 
         </mn> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           2 
         </mn> 
        </msub> 
       </mrow> 
       <mo>
         } 
       </mo> 
      </mrow> 
     </mrow> 
    </math></p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         1 
       </mn> 
      </msub> 
     </mrow> 
    </math> represents the first coefficient influencing the model’s predictions, and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         2 
       </mn> 
      </msub> 
     </mrow> 
    </math> represents the second coefficient. The parameters in 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       Θ 
     </mi> 
    </math> are crucial for minimizing 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        J 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           1 
         </mn> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           2 
         </mn> 
        </msub> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> while adhering to the constraints that require both 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         1 
       </mn> 
      </msub> 
     </mrow> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         2 
       </mn> 
      </msub> 
     </mrow> 
    </math> to be non-negative. This ensures that the model parameters remain within a physically meaningful range, ultimately leading to a more accurate representation of the relationship between the inputs and outputs.</p>
   <p>Mathematically, the objective function can be expressed as:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        min 
      </mi> 
      <mi>
        J 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           1 
         </mn> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           2 
         </mn> 
        </msub> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <msubsup> 
       <mrow> 
        <mrow> 
         <mo>
           ‖ 
         </mo> 
         <mrow> 
          <msub> 
           <mi>
             X 
           </mi> 
           <mrow> 
            <mtext>
              train 
            </mtext> 
           </mrow> 
          </msub> 
          <mo>
            − 
          </mo> 
          <mi>
            Y 
          </mi> 
         </mrow> 
         <mo>
           ‖ 
         </mo> 
        </mrow> 
       </mrow> 
       <mi>
         F 
       </mi> 
       <mn>
         2 
       </mn> 
      </msubsup> 
      <mo>
        + 
      </mo> 
      <mrow> 
       <mo>
         { 
       </mo> 
       <mrow> 
        <mtable columnalign="left"> 
         <mtr columnalign="left"> 
          <mtd columnalign="left"> 
           <mrow> 
            <mn>
              1 
            </mn> 
            <mo>
              × 
            </mo> 
            <msup> 
             <mrow> 
              <mn>
                10 
              </mn> 
             </mrow> 
             <mn>
               6 
             </mn> 
            </msup> 
           </mrow> 
          </mtd> 
          <mtd columnalign="left"> 
           <mrow> 
            <mtext>
              if 
            </mtext> 
            <mtext>
                
            </mtext> 
            <msub> 
             <mi>
               θ 
             </mi> 
             <mn>
               1 
             </mn> 
            </msub> 
            <mo>
              &lt; 
            </mo> 
            <mn>
              0 
            </mn> 
            <mtext>
                
            </mtext> 
            <mtext>
              or 
            </mtext> 
            <mtext>
                
            </mtext> 
            <msub> 
             <mi>
               θ 
             </mi> 
             <mn>
               2 
             </mn> 
            </msub> 
            <mo>
              &lt; 
            </mo> 
            <mn>
              0 
            </mn> 
           </mrow> 
          </mtd> 
         </mtr> 
         <mtr columnalign="left"> 
          <mtd columnalign="left"> 
           <mn>
             0 
           </mn> 
          </mtd> 
          <mtd columnalign="left"> 
           <mrow> 
            <mtext>
              otherwise 
            </mtext> 
           </mrow> 
          </mtd> 
         </mtr> 
        </mtable> 
       </mrow> 
      </mrow> 
      <mo>
        + 
      </mo> 
      <mi>
        λ 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msubsup> 
         <mi>
           θ 
         </mi> 
         <mn>
           1 
         </mn> 
         <mn>
           2 
         </mn> 
        </msubsup> 
        <mo>
          + 
        </mo> 
        <msubsup> 
         <mi>
           θ 
         </mi> 
         <mn>
           2 
         </mn> 
         <mn>
           2 
         </mn> 
        </msubsup> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math>(10)</p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        J 
      </mi> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           1 
         </mn> 
        </msub> 
        <mo>
          , 
        </mo> 
        <msub> 
         <mi>
           θ 
         </mi> 
         <mn>
           2 
         </mn> 
        </msub> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
     </mrow> 
    </math> measures how well the model’s predictions (calculated outflows) match the actual observed outflows. The goal is to find the parameters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         1 
       </mn> 
      </msub> 
     </mrow> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         2 
       </mn> 
      </msub> 
     </mrow> 
    </math> that minimize this function. The Frobenius norm 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msubsup> 
       <mrow> 
        <mrow> 
         <mo>
           ‖ 
         </mo> 
         <mrow> 
          <mtext>
              
          </mtext> 
          <mo>
            . 
          </mo> 
          <mtext>
              
          </mtext> 
         </mrow> 
         <mo>
           ‖ 
         </mo> 
        </mrow> 
       </mrow> 
       <mi>
         F 
       </mi> 
       <mn>
         2 
       </mn> 
      </msubsup> 
     </mrow> 
    </math> quantifies the difference between the predicted and actual values. Here, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         X 
       </mi> 
       <mrow> 
        <mtext>
          train 
        </mtext> 
       </mrow> 
      </msub> 
     </mrow> 
    </math> represents the observed outflows, while 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       Y 
     </mi> 
    </math> denotes the calculated outflows based on the model and parameters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         1 
       </mn> 
      </msub> 
     </mrow> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         2 
       </mn> 
      </msub> 
     </mrow> 
    </math>. The calculated outflows 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       Y 
     </mi> 
    </math> can be expressed as:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        Y 
      </mi> 
      <mo>
        = 
      </mo> 
      <mi>
        W 
      </mi> 
      <mo>
        ⋅ 
      </mo> 
      <mi>
        H 
      </mi> 
     </mrow> 
    </math>(11)</p>
   <p>where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       W 
     </mi> 
    </math> is the matrix of features and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       H 
     </mi> 
    </math> is the matrix of coefficients derived from the parameters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         1 
       </mn> 
      </msub> 
     </mrow> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         θ 
       </mi> 
       <mn>
         2 
       </mn> 
      </msub> 
     </mrow> 
    </math>. The Frobenius norm computes the squared difference between these two sets of values, aggregating the errors across all observations. To discourage invalid parameter values, a penalty term is added, ensuring that only non-negative parameters are considered in the optimization process. To prevent overfitting by penalizing large values of the parameters, we add a regularization coefficient 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       λ 
     </mi> 
    </math> that controls the strength of this penalty, balancing the fit of the model to the data against the complexity of the model.</p>
   <p>Document Data Sets:</p>
   <p>We cluster the 14 document data sets from <xref ref-type="bibr" rid="scirp.146377-38">
     [38]
    </xref> using KL-ONMF. <xref ref-type="table" rid="table1">
     Table 1
    </xref> provides not only the names of the data sets but also their dimensions (where 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       m 
     </mi> 
    </math> represents the number of words and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       n 
     </mi> 
    </math> the number of documents) and the number of clusters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       r 
     </mi> 
    </math>.</p>
   <table-wrap id="table1">
    <label>
     <xref ref-type="table" rid="table1">
      Table 1
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146377-"></xref>Table 1. Summary of text datasets (for each dataset, 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  n
 
       </mi>

      </math> is the total number of documents, 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  m
 
       </mi>

      </math> the total number of words, 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        
  r
 
       </mi>

      </math> the number of classes, 

      <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
  
        <msub> 
   
         <mover accent="true"> 
    
          <mi>
           
     n
    
          </mi> 
    
          <mo>
           
     ¯
    
          </mo> 
   
         </mover> 
   
         <mi>
          
    c
   
         </mi> 
  
        </msub> 
 
       </mrow>

      </math> the average number of documents per class, and Balance the size ratio of the smallest class to the largest class).</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Data</p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           n 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           m 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           r 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
          <msub> 
           <mover accent="true"> 
            <mi>
              n 
            </mi> 
            <mo>
              ¯ 
            </mo> 
           </mover> 
           <mi>
             c 
           </mi> 
          </msub> 
         </mrow> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Balance</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">ngsim</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">2998</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">15,810</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">3</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center"></p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center"></p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">classic</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1073</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">30</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">20</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center"></p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">ohscal</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">11,162</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">11,465</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1116</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.437</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">k1b</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">2340</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">21,839</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">390</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.043</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">hitech</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">2301</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">10,080</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">384</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.192</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">reviews</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">4069</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">18,483</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">5</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">814</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.098</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">sports</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">8580</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">14,870</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">7</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1226</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.036</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">la1</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">3204</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">534</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.290</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">la12</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6279</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1047</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.282</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">la2</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">3075</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">513</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.274</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr11</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">414</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6429</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">9</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">46</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.046</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr23</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">204</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">5832</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">34</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.066</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr41</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">878</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">7454</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">88</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.037</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr45</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">690</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">8261</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">69</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.088</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>Accuracy:</p>
   <p>Accuracy is defined as the proportion of correctly classified data points relative to the total number of data points. In the setting of document clustering, the accuracy of a computed disjoint clustering 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msubsup> 
       <mrow> 
        <mrow> 
         <mo>
           { 
         </mo> 
         <mrow> 
          <msub> 
           <mover accent="true"> 
            <mi>
              C 
            </mi> 
            <mo>
              ˜ 
            </mo> 
           </mover> 
           <mi>
             i 
           </mi> 
          </msub> 
         </mrow> 
         <mo>
           } 
         </mo> 
        </mrow> 
       </mrow> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          = 
        </mo> 
        <mn>
          1 
        </mn> 
       </mrow> 
       <mi>
         r 
       </mi> 
      </msubsup> 
     </mrow> 
    </math> compared to the true disjoint clusters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msub> 
       <mi>
         C 
       </mi> 
       <mi>
         i 
       </mi> 
      </msub> 
      <mo>
        ⊂ 
      </mo> 
      <mrow> 
       <mo>
         { 
       </mo> 
       <mrow> 
        <mn>
          1 
        </mn> 
        <mo>
          , 
        </mo> 
        <mn>
          2 
        </mn> 
        <mo>
          , 
        </mo> 
        <mo>
          ⋯ 
        </mo> 
        <mo>
          , 
        </mo> 
        <mi>
          n 
        </mi> 
       </mrow> 
       <mo>
         } 
       </mo> 
      </mrow> 
     </mrow> 
    </math> for 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mn>
        1 
      </mn> 
      <mo>
        ≤ 
      </mo> 
      <mi>
        i 
      </mi> 
      <mo>
        ≤ 
      </mo> 
      <mi>
        r 
      </mi> 
     </mrow> 
    </math> is defined as follows:</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mtext>
        accuracy 
      </mtext> 
      <mrow> 
       <mo>
         ( 
       </mo> 
       <mrow> 
        <msubsup> 
         <mrow> 
          <mrow> 
           <mo>
             { 
           </mo> 
           <mrow> 
            <msub> 
             <mover accent="true"> 
              <mi>
                C 
              </mi> 
              <mo>
                ˜ 
              </mo> 
             </mover> 
             <mi>
               i 
             </mi> 
            </msub> 
           </mrow> 
           <mo>
             } 
           </mo> 
          </mrow> 
         </mrow> 
         <mrow> 
          <mi>
            i 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           r 
         </mi> 
        </msubsup> 
       </mrow> 
       <mo>
         ) 
       </mo> 
      </mrow> 
      <mo>
        = 
      </mo> 
      <munder> 
       <mrow> 
        <mtext>
          min 
        </mtext> 
       </mrow> 
       <mrow> 
        <mi>
          π 
        </mi> 
        <mo>
          ∈ 
        </mo> 
        <mrow> 
         <mo>
           [ 
         </mo> 
         <mrow> 
          <mn>
            1 
          </mn> 
          <mo>
            , 
          </mo> 
          <mn>
            2 
          </mn> 
          <mo>
            , 
          </mo> 
          <mo>
            ⋯ 
          </mo> 
          <mo>
            , 
          </mo> 
          <mi>
            r 
          </mi> 
         </mrow> 
         <mo>
           ] 
         </mo> 
        </mrow> 
       </mrow> 
      </munder> 
      <mfrac> 
       <mn>
         1 
       </mn> 
       <mi>
         n 
       </mi> 
      </mfrac> 
      <munderover> 
       <mstyle displaystyle="true" mathsize="140%"> 
        <mo>
          ∑ 
        </mo> 
       </mstyle> 
       <mrow> 
        <mi>
          i 
        </mi> 
        <mo>
          = 
        </mo> 
        <mn>
          1 
        </mn> 
       </mrow> 
       <mi>
         r 
       </mi> 
      </munderover> 
      <mrow> 
       <mo>
         | 
       </mo> 
       <mrow> 
        <msub> 
         <mi>
           C 
         </mi> 
         <mi>
           i 
         </mi> 
        </msub> 
        <mo>
          ∩ 
        </mo> 
        <msub> 
         <mover accent="true"> 
          <mi>
            C 
          </mi> 
          <mo>
            ˜ 
          </mo> 
         </mover> 
         <mrow> 
          <mi>
            π 
          </mi> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mi>
             i 
           </mi> 
           <mo>
             ) 
           </mo> 
          </mrow> 
         </mrow> 
        </msub> 
       </mrow> 
       <mo>
         | 
       </mo> 
      </mrow> 
      <mo>
        . 
      </mo> 
     </mrow> 
    </math>(12)</p>
   <p>The correctly classified data points are those that have been assigned to the same cluster as their original data. The clustering process was carried out using the K-means algorithm.</p>
   <p>Running time:</p>
   <p>
    <xref ref-type="table" rid="table2">
     Table 2
    </xref> summarizes the execution times associated with the objective function evaluated through the Nelder-Mead optimization algorithm across various datasets.</p>
   <table-wrap id="table2">
    <label>
     <xref ref-type="table" rid="table2">
      Table 2
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146377-"></xref>Table 2. Execution times of datasets.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Dataset</p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Min Time (s)</p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Max Time (s)</p></td> 
      <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Mean Time (s)</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">ng3sim</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">1.40</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">2.57</p></td> 
      <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">1.84</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">classic</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">172.11</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">215.82</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">193.40</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">ohscal</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">4.45</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">31.01</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">13.46</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">k1b</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">30.78</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">31.28</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">31.05</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">hitech</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.58</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.80</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.65</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">reviews</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">4.60</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">48.19</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">19.15</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">sports</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">3.96</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6.90</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">5.48</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">la1</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">2.56</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">2.78</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">2.66</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">la12</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">4.39</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">8.58</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">6.30</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">la2</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">2.12</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">55.91</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">20.07</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr11</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.91</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.18</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.01</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr23</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.73</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.75</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">0.74</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr41</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.14</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.50</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.31</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="22.61%"><p style="text-align:center">tr45</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.53</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.67</p></td> 
      <td class="acenter" width="22.61%"><p style="text-align:center">1.60</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>The results indicate substantial variability in processing times, with the “classic” dataset exhibiting the highest execution times, ranging from 172.11 to 215.82 seconds. This suggests that this dataset may present more complex optimization challenges or larger data sizes. Conversely, datasets like “ng3sim” and “hitech” show significantly lower mean execution times of 1.84 and 1.65 seconds, respectively, indicating efficient convergence during optimization. The “la2” dataset, with a maximum time of 55.91 seconds, points to potential outliers that may require further analysis to understand the underlying causes of increased processing time. Meanwhile, datasets such as “tr11,” “tr23,” “tr41,” and “tr45” demonstrate consistent execution times, which may reflect similar optimization landscapes. Overall, these execution times underscore the impact of dataset characteristics on the performance of the Nelder-Mead optimization algorithm, highlighting opportunities for further refinement in handling more computationally intensive datasets.</p>
   <p>Optimal Parameters</p>
   <p>In this section, we present the optimal parameters identified for each dataset through cross-validation. The table below summarizes these parameters, which play a crucial role in enhancing the performance of our model. The first column of the table lists the datasets. To ensure stability and prevent excessive fluctuations during optimization, we fixed the parameter 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> at 0.0001 rather than allowing it to be optimized alongside 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math>. This choice is guided by the observation that large variations in 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> can lead to significant distortions in the TF-IDF scores, potentially overshadowing the contributions of the other parameters. By maintaining 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> at a low constant value, we provide a baseline influence that stabilizes the model while allowing 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> to be fine-tuned for better performance.</p>
   <p>In a sensitivity analysis, we examined the effects of varying 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> around its fixed value. The results indicated that small deviations from 0.0001 did not substantially alter the overall performance of the clustering algorithms. However, larger adjustments (e.g., setting 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> to 0.01 or higher) resulted in noticeable degradation in clustering quality, highlighting the importance of keeping 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> within a narrow range. This analysis supports our decision to fix 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> at 0.0001, ensuring that the model remains robust while still allowing 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> the flexibility needed to adapt to specific dataset characteristics. The remaining columns represent the optimized parameters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math>, which vary for each dataset. The parameter 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> influences the model’s responsiveness to changes in the data, while 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> affects the overall scaling of the model’s output. Finally, the “Best Smooth” column indicates the optimal smoothing value achieved during cross-validation, which reflects the model’s performance in terms of clustering accuracy. The varying values in this column highlight how different datasets require tailored parameter settings to achieve the best results.</p>
   <table-wrap id="table3">
    <label>
     <xref ref-type="table" rid="table3">
      Table 3
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146377-"></xref>Table 3. Optimal parameters for transformations.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="9.99%"><p style="text-align:center">Dataset</p></td> 
      <td class="custom-bottom-td acenter" width="29.99%" colspan="3"><p style="text-align:center">Log</p></td> 
      <td class="custom-bottom-td acenter" width="30.01%" colspan="3"><p style="text-align:center">Sq</p></td> 
      <td class="custom-bottom-td acenter" width="30.01%" colspan="3"><p style="text-align:center">HTan</p></td> 
     </tr> 
     <tr> 
      <td class="custom-bottom-td custom-top-td acenter" width="9.99%"><p style="text-align:center"></p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="9.99%"><p style="text-align:center">b</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="9.99%"><p style="text-align:center">c</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">BS</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">b</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">c</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">BS</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">b</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">c</p></td> 
      <td class="custom-bottom-td custom-top-td acenter" width="10.00%"><p style="text-align:center">BS</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="9.99%"><p style="text-align:center">ng3sim</p></td> 
      <td class="custom-top-td acenter" width="9.99%"><p style="text-align:center">2.7809</p></td> 
      <td class="custom-top-td acenter" width="9.99%"><p style="text-align:center">9.1445</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">50.1700</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">2.7951</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">9.1290</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">7.0700</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">2.3249</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">9.1467</p></td> 
      <td class="custom-top-td acenter" width="10.00%"><p style="text-align:center">18.3700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">classic</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">0.0682</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">1.0919</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">25.0000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.0682</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">1.0919</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">15.0000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.0682</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">1.0918</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">15.0000</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">ohscal</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">3.0457</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">16.4311</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.1700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.4546</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">16.5128</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">25.0000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.6363</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">16.3386</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">17.0300</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">k1b</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.2467</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">21.8494</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">16.7300</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">1.6804</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">22.9310</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.1000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.7789</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">22.8170</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.8700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">hitech</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">3.1803</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">23.7720</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">34.0000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.1843</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">23.8005</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">8.5700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.6590</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">23.7885</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.2000</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">reviews</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">3.0342</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">27.0628</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">66.6700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.0326</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">27.0573</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">13.3700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.0317</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">27.0648</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">83.3300</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">sports</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.5719</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">20.4878</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">8.5000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.5922</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">20.4682</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">35.0000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.0420</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">20.4616</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">66.6700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">la1</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.7167</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">17.7621</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.4000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.2277</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">17.6894</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">46.6700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.2066</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">17.8179</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">7.1700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">la12</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.2671</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">18.6777</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">16.7000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.7673</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">18.6036</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">8.3700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.7632</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">18.6087</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">41.6700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">la2</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.8408</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">17.7540</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">23.3400</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.3745</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">17.6619</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">1.7000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.8439</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">17.7351</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.2300</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">tr11</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.4381</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">5.7950</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">50.6700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">1.2637</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">5.3720</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">40.0300</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.4959</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">6.1245</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">8.3700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">tr23</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">1.8971</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">3.3695</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.1000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.2632</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.5174</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.1000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.2537</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">3.5558</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.1000</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">tr41</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.5543</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">16.7309</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">13.3700</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.2009</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">16.4067</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.7300</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.1465</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">15.7258</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">18.3700</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="9.99%"><p style="text-align:center">tr45</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">2.3656</p></td> 
      <td class="acenter" width="9.99%"><p style="text-align:center">9.2114</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">16.7300</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">2.3873</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">8.8119</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">0.1000</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">1.5483</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">8.2574</p></td> 
      <td class="acenter" width="10.00%"><p style="text-align:center">17.6700</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>
    <xref ref-type="table" rid="table3">
     Table 3
    </xref> presents parameter values for three different nonlinear transformation: logarithmic transformation (Equation (7)), square root transformation (Equation (8)), and hyperbolic tangent transformation (Equation (9)). Even though these methods use different mathematical functions, the values for parameters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> are very similar across all datasets. For the logarithmic transformation, 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> usually ranges from 0.068 to 3.180, and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> ranges from 1.091 to 27.062. The square root transformation also shows similar patterns, with 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> values in the same range and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> values that suggest these two transformations might work well with similar types of data. By analyzing the ranges of the parameter of the “Best Smooth”, we observe interesting trends: for the logarithmic transformation, the values of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> range from 0.068 to 3.180, while those for 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> span from 1.091 to 27.062. The “Best Smooth” values in this transformation vary from 0.10 to 66.67. For the square root transformation, the 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> parameters remain within a similar range, from 0.068 to 3.374, and the values of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> range from 1.091 to 27.057. The “Best Smooth” values for this function is between 0.10 and 46.67. For the hyperbolic tangent transformation, the 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> values also fall within the same interval, from 0.068 to 3.636, with 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> values ranging from 1.091 to 27.064. The “Best Smooth” values in this approach vary from 0.10 to 83.33. These intervals demonstrate that, although the transformations employ different mathematical nonlinear functions, the parameters 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> exhibit significant similarities, suggesting that these transformations may be well-suited for similar types of data. The choice of a fixed starting value for 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math> allows for a clearer analysis of the impact of 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       b 
     </mi> 
    </math> and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       c 
     </mi> 
    </math> on the results, minimizing the effects of variations in 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
       a 
     </mi> 
    </math>.</p>
   <p>Performance Evaluation:</p>
   <p>We present the accuracy achieved by KL-ONMF on both the original and transformed data across different datasets. The accuracy metrics are denoted as follows: acc-KL for overall accuracy, BestAcKL-ONMF (Train) for the accuracy on the training dataset, and BestAcKL-ONMF (Validation) for the accuracy on the validation dataset. The distinction between “training” and “validation” phases is critical in our approach, particularly given the nature of the KL-ONMF method and the transformations applied to the data. During the training phase, the model focuses on learning complex patterns and relationships within the feature space that is generated from the transformed TF-IDF scores. This phase heavily relies on the training data, which plays a significant role in adjusting the model’s parameters. Conversely, the validation phase serves an essential purpose by preventing the model from overfitting to the training data. Evaluating the model on a separate validation set allows us to gauge its ability to generalize to new, unseen data—an aspect that is vital for practical applications. In our methodology, where TF-IDF scores are subjected to various transformations, validating performance on data that the model has not previously encountered becomes particularly important. This validation helps us assess the effectiveness of the smoothing and non-linear transformations applied, ensuring they genuinely enhance the model’s clustering and classification capabilities rather than simply memorizing the training data. Thus, the clear separation of training and validation phases in our approach not only bolsters the reliability of the model’s performance but also underscores its robustness in real-world scenarios.</p>
   <table-wrap id="table4">
    <label>
     <xref ref-type="table" rid="table4">
      Table 4
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146377-"></xref>Table 4. KL-ONMF vs. KL-ONMF Transformed for the clustering of 14 document data sets: Cross-Validation with Repetitions for Logarithm Transformation.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="14.28%"><p style="text-align:center">Dataset</p></td> 
      <td class="custom-bottom-td acenter" width="10.44%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           m 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="10.44%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           n 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="7.50%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           r 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="14.71%"><p style="text-align:center">AcKL-ONMF</p></td> 
      <td class="custom-bottom-td acenter" width="19.11%"><p style="text-align:center">Best AcKL-ONMF (Train)</p></td> 
      <td class="custom-bottom-td acenter" width="23.51%"><p style="text-align:center">Best AcKL-ONMF (Validation)</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="14.28%"><p style="text-align:center">ng3sim</p></td> 
      <td class="custom-top-td acenter" width="10.44%"><p style="text-align:center">2998</p></td> 
      <td class="custom-top-td acenter" width="10.44%"><p style="text-align:center">15,810</p></td> 
      <td class="custom-top-td acenter" width="7.50%"><p style="text-align:center">3</p></td> 
      <td class="custom-top-td acenter" width="14.71%"><p style="text-align:center">71.45</p></td> 
      <td class="custom-top-td acenter" width="19.11%"><p style="text-align:center">74.61</p></td> 
      <td class="custom-top-td acenter" width="23.51%"><p style="text-align:center">74.12</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">classic</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">7094</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">41,681</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">4</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">85.38</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">91.99</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">85.49</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">ohscal</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">11,162</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">11,465</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">47.74</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">51.69</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">51.20</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">k1b</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">2340</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">21,819</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">58.46</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">60.94</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">60.53</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">hitech</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">2301</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">10,080</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">38.55</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">37.81</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">37.27</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">reviews</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">4069</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">18,483</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">5</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">72.35</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">72.12</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">71.84</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">sports</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">8580</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">14,870</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">7</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">66.55</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">67.19</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">65.45</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la1</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">3204</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">61.20</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">63.10</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">61.66</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la12</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">6279</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">67.48</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">63.10</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">61.75</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la2</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">3075</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">59.90</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">63.91</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">61.95</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr11</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">414</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">6424</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">9</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">54.11</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">47.83</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">44.85</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr23</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">204</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">5831</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">34.31</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">43.14</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">41.83</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr41</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">878</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">7453</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">48.63</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">50.72</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">49.24</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr45</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">690</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center">8261</p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">59.57</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">60.39</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">58.89</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">Averages</p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="10.44%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="7.50%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="14.71%"><p style="text-align:center">58.98</p></td> 
      <td class="acenter" width="19.11%"><p style="text-align:center">60.61</p></td> 
      <td class="acenter" width="23.51%"><p style="text-align:center">59.01</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>
    <xref ref-type="table" rid="table4">
     Table 4
    </xref> compares the performance of the AcKL-ONMF method applied to both untransformed data and smoothed, non-linear transformations. The metrics include baseline accuracy, Best AcKL-ONMF accuracy for training, and validation sets across various datasets.</p>
   <p>Notably, the Best AcKL-ONMF(Train) values consistently outperform the baseline accuracy, indicating that the method effectively learns features from the transformed data. For instance, Dataset 1 shows a significant increase from 71.45% to 74.61%, demonstrating enhanced model performance through the application of smoothing and non-linear transformations.</p>
   <p>When examining the Best AcKL-ONMF(Validation) results, there is a clear trend of improvements in accuracy on unseen data. Datasets such as Dataset 2 show a rise from 85.38% to 91.99% in training, with a validation accuracy of 85.49%, illustrating the model’s ability to generalize effectively.</p>
   <p>The average accuracies further validate our approach. The training set achieves an average of 60.61%, while the validation set maintains competitive performance at 59.01%. This consistency across both training and validation sets suggests that the AcKL-ONMF method not only optimizes the model parameters well but also ensures robust performance on unseen data.</p>
   <table-wrap id="table5">
    <label>
     <xref ref-type="table" rid="table5">
      Table 5
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146377-"></xref>Table 5. KL-ONMF vs. KL-ONMF Transformed for the clustering of 14 document data sets: Cross-Validation with Repetitions for Squared Transformation.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="14.28%"><p style="text-align:center">Dataset</p></td> 
      <td class="custom-bottom-td acenter" width="11.91%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           m 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="11.91%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           n 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="11.91%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           r 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="16.65%"><p style="text-align:center">AcKL-ONMF</p></td> 
      <td class="custom-bottom-td acenter" width="16.66%"><p style="text-align:center">Best AcKL-ONMF (Train)</p></td> 
      <td class="custom-bottom-td acenter" width="16.66%"><p style="text-align:center">Best AcKL-ONMF (Validation)</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="14.28%"><p style="text-align:center">ng3sim</p></td> 
      <td class="custom-top-td acenter" width="11.91%"><p style="text-align:center">2998</p></td> 
      <td class="custom-top-td acenter" width="11.91%"><p style="text-align:center">15,810</p></td> 
      <td class="custom-top-td acenter" width="11.91%"><p style="text-align:center">3</p></td> 
      <td class="custom-top-td acenter" width="16.65%"><p style="text-align:center">71.45</p></td> 
      <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">73.85</p></td> 
      <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">72.98</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">classic</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">7094</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">41,681</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">4</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">85.38</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">92.92</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">92.46</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">ohscal</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">11,162</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">11,465</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">47.74</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">51.74</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">51.18</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">k1b</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">2340</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">21,819</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">58.46</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">56.17</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">55.98</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">hitech</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">2301</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10,080</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">38.55</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">38.06</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">37.87</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">reviews</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">4069</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">18,483</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">5</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">72.35</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">69.14</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">68.15</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">sports</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">8580</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">14,870</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">7</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">66.55</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">68.56</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">68.24</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la1</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">3204</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">61.20</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">61.49</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">60.21</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la12</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6279</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">67.48</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">60.05</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">57.47</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la2</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">3075</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">59.90</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">62.12</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">60.11</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr11</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">414</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6424</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">9</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">54.11</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">45.65</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">44.85</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr23</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">204</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">5831</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">34.31</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">40.52</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">41.01</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr41</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">878</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">7453</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">48.63</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">50.30</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">48.71</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr45</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">690</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">8261</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">59.57</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">60.77</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">59.47</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">Averages</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">58.98</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">59.38</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">58.48</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>
    <xref ref-type="table" rid="table5">
     Table 5
    </xref> compares the performance of the AcKL-ONMF method applied to both untransformed data and smoothed, non-linear transformations. The metrics include baseline accuracy, Best AcKL-ONMF accuracy for training, and validation sets across various datasets.</p>
   <p>Notably, the Best AcKL-ONMF(Train) values consistently outperform the baseline accuracy, indicating that the method effectively learns features from the transformed data. For instance, Dataset 1 shows a significant increase from 71.45% to 73.85%, demonstrating enhanced model performance through the application of smoothing and non-linear transformations.</p>
   <p>When examining the Best AcKL-ONMF(Validation) results, there is a clear trend of improvements in accuracy on unseen data. Datasets such as Dataset 2 show a rise from 85.38% to 92.92% in training, with a validation accuracy of 92.46%, illustrating the model’s ability to generalize effectively.</p>
   <p>The average accuracies further validate our approach. The training set achieves an average of 59.38%, while the validation set maintains competitive performance at 58.48%. This consistency across both training and validation sets suggests that the AcKL-ONMF method not only optimizes the model parameters well but also ensures robust performance on unseen data.</p>
   <table-wrap id="table6">
    <label>
     <xref ref-type="table" rid="table6">
      Table 6
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146377-"></xref>Table 6. KL-ONMF vs. KL-ONMF Transformed for the clustering of 14 document data sets: Cross-Validation with Repetitions for Hyperbolic Tangent Transformation.</title>
    </caption>
    <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
     <tr> 
      <td class="custom-bottom-td acenter" width="14.28%"><p style="text-align:center">Dataset</p></td> 
      <td class="custom-bottom-td acenter" width="11.91%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           m 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="11.91%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           n 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="11.91%"><p style="text-align:center"> 
        <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
           r 
         </mi> 
        </math></p></td> 
      <td class="custom-bottom-td acenter" width="16.65%"><p style="text-align:center">AcKL-ONMF</p></td> 
      <td class="custom-bottom-td acenter" width="16.66%"><p style="text-align:center">Best AcKL-ONMF (Train)</p></td> 
      <td class="custom-bottom-td acenter" width="16.66%"><p style="text-align:center">Best AcKL-ONMF (Validation)</p></td> 
     </tr> 
     <tr> 
      <td class="custom-top-td acenter" width="14.28%"><p style="text-align:center">ng3sim</p></td> 
      <td class="custom-top-td acenter" width="11.91%"><p style="text-align:center">2998</p></td> 
      <td class="custom-top-td acenter" width="11.91%"><p style="text-align:center">15,810</p></td> 
      <td class="custom-top-td acenter" width="11.91%"><p style="text-align:center">3</p></td> 
      <td class="custom-top-td acenter" width="16.65%"><p style="text-align:center">71.45</p></td> 
      <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">73.86</p></td> 
      <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">72.95</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">classic</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">7094</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">41,681</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">4</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">85.38</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">91.58</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">82.24</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">ohscal</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">11,162</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">11,465</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">47.74</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">51.57</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">51.15</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">k1b</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">2340</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">21,819</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">58.46</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">60.23</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">59.46</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">hitech</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">2301</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10,080</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">38.55</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">38.78</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">38.17</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">reviews</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">4069</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">18,483</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">5</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">72.35</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">71.27</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">71.15</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">sports</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">8580</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">14,870</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">7</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">66.55</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">68.67</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">67.88</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la1</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">3204</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">61.20</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">66.46</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">63.77</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la12</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6279</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">67.48</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">61.03</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">57.76</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">la2</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">3075</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">31,472</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">59.90</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">63.12</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">62.35</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr11</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">414</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6424</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">9</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">54.11</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">50.32</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">46.62</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr23</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">204</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">5831</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">6</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">34.31</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">45.10</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">39.54</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr41</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">878</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">7453</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">48.63</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">51.78</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">49.77</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">tr45</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">690</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">8261</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center">10</p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">59.57</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">65.75</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">64.06</p></td> 
     </tr> 
     <tr> 
      <td class="acenter" width="14.28%"><p style="text-align:center">Averages</p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="11.91%"><p style="text-align:center"></p></td> 
      <td class="acenter" width="16.65%"><p style="text-align:center">58.98</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">61.39</p></td> 
      <td class="acenter" width="16.66%"><p style="text-align:center">59.06</p></td> 
     </tr> 
    </table>
   </table-wrap>
   <p>
    <xref ref-type="table" rid="table6">
     Table 6
    </xref> compares the performance of the AcKL-ONMF method applied to both untransformed and smoothed, non-linear transformed data. The metrics include baseline accuracy, Best AcKL-ONMF accuracy for training, and validation sets across various datasets.</p>
   <p>The Best AcKL-ONMF(Train) values consistently exceed the baseline accuracy, demonstrating the method’s effectiveness in extracting features from the transformed data. For example, Dataset 1 shows an increase from 71.45% to 73.86%, indicating a substantial improvement in model performance due to the application of smoothing and non-linear transformations.</p>
   <p>In terms of Best AcKL-ONMF(Validation), the results reveal a positive trend in accuracy for unseen data. Dataset 2 exemplifies this, with the accuracy rising from 85.38% to 91.58% in training, while validation accuracy stands at 82.24%. This illustrates the model’s capacity to generalize effectively to new data.</p>
   <p>Average accuracies further support our findings, with the training set achieving an average of 61.39%, compared to a baseline of 58.98%. The validation set maintains a competitive average of 59.06%, reinforcing the consistency of the AcKL-ONMF method across both training and validation contexts.</p>
  </sec><sec id="s4">
   <title>5. Conclusion and Suggestion</title>
   <p>In this paper, we developed an innovative clustering approach for non-negative data, named Optimal Predictive Modeling of Nonlinear Transformations using the Kullback-Leibler Divergence “OPMNT-KL-ONMF”. Our methodology begins by applying smoothing to TF-IDF scores, followed by the integration of parameters optimized using the Nelder-Mead method. We then employ nonlinear transformations to extract nonlinear features. Comparative results between KL-ONMF applied to original data and data transformed by OPMNT demonstrate a significant performance improvement. Our OPMNT-KL-ONMF model represents a promising advancement in the field of clustering non-negative data, such as document datasets, and paves the way for future research and applications across various domains of data analysis.</p>
  </sec><sec id="s5">
   <title>Acknowledgements</title>
   <p>Sincere thanks to the members of OJAppS for their exemplary professionalism, and a special acknowledgment to Managing Editor Emily Lee for her exceptional commitment to high-quality standards.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.146377-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Paatero, P. and Tapper, U. (1994) Positive Matrix Factorization: A Non-Negative Factor Model with Optimal Utilization of Error Estimates of Data Values. Environmetrics, 5, 111-126. &gt;https://doi.org/10.1002/env.3170050203 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     You, M., Wang, H., Liu, Z., Chen, C., Liu, J., Xu, X., et al. (2017) Novel Feature Extraction Method for Cough Detection Using NMF. IET Signal Processing, 11, 515-520. &gt;https://doi.org/10.1049/iet-spr.2016.0341 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Oti, E.U., Olusola, M.O., Eze, F.C. and Enogwe, S.U. (2021) Comprehensive Review of K-Means Clustering Algorithms. International Journal of Advances in Scientific Research and Engineering, 7, 64-69. &gt;https://doi.org/10.31695/ijasre.2021.34050 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Saberi-Movahed, F., Berahman, K., Sheikhpour, R., Li, Y. and Pan, S. (2024) Nonnegative Matrix Factorization in Dimensionality Reduction: A Survey. 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, Y., Jia, Y., Hu, C. and Turk, M. (2005) Non-Negative Matrix Factorization Framework for Face Recognition. International Journal of Pattern Recognition and Artificial Intelligence, 19, 495-511. &gt;https://doi.org/10.1142/s0218001405004198
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mohammadiha, N., Smaragdis, P. and Leijon, A. (2013) Supervised and Unsupervised Speech Enhancement Using Nonnegative Matrix Factorization. IEEE Transactions on Audio, Speech, and Language Processing, 21, 2140-2151. &gt;https://doi.org/10.1109/tasl.2013.2270369 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wilson, K.W., Raj, B., Smaragdis, P. and Divakaran, A. (2008) Speech Denoising Using Nonnegative Matrix Factorization with Priors. 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, Las Vegas, 31 March-4 April 2008, 4029-4032. &gt;https://doi.org/10.1109/icassp.2008.4518538 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pauca, V.P., Piper, J. and Plemmons, R.J. (2006) Nonnegative Matrix Factorization for Spectral Data Analysis. Linear Algebra and Its Applications, 416, 29-47. &gt;https://doi.org/10.1016/j.laa.2005.06.025 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Shiga, M. and Muto, S. (2019) Non-Negative Matrix Factorization and Its Extensions for Spectral Image Data Analysis. e-Journal of Surface Science and Nanotechnology, 17, 148-154. &gt;https://doi.org/10.1380/ejssnt.2019.148 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Devarajan, K. (2008) Nonnegative Matrix Factorization: An Analytical and Interpretive Tool in Computational Biology. PLOS Computational Biology, 4, e1000029. &gt;https://doi.org/10.1371/journal.pcbi.1000029 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Qi, Q., Zhao, Y., Li, M. and Simon, R. (2009) Non-Negative Matrix Factorization of Gene Expression Profiles: A Plug-In for Brb-Arraytools. Bioinformatics, 25, 545-547. &gt;https://doi.org/10.1093/bioinformatics/btp009 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ouyang, D., Miao, R., Zeng, J., Li, X., Ai, N., Wang, P., et al. (2024) Splhrnmtf: Robust Orthogonal Non-Negative Matrix Tri-Factorization with Self-Paced Learning and Dual Hypergraph Regularization for Predicting miRNA-Disease Associations. BMC Genomics, 25, Article No. 885. &gt;https://doi.org/10.1186/s12864-024-10729-w 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Li, L., Gao, Z., Wang, Y.-T., Zhang, M.-W., Ni, J.-C., Zheng, C.-H. and Su, Y. (2021) SCMFDA: Predicting MicroRNA-Disease Associations Based on Similarity Constrained Matrix Factorization. PLOS Computational Biology, 17, e1009165.
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gligorijevic, V., Panagakis, Y. and Zafeiriou, S. (2018) Non-Negative Matrix Factorizations for Multiplex Network Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41, 928-940. &gt;https://doi.org/10.1109/tpami.2018.2821146
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhang, Z.-Y. (2012) Nonnegative Matrix Factorization: Models, Algorithms and Applications. In: Intelligent Systems Reference Library, Springer, 99-134. &gt;https://doi.org/10.1007/978-3-642-23241-1_6
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jia, X. and Guo, B. (2022) Non-Negative Matrix Factorization Based on Smoothing and Sparse Constraints for Hyperspectral Unmixing. Sensors, 22, Article 5417. &gt;https://doi.org/10.3390/s22145417 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jiménez-Sánchez, D., Ariz, M., Morgado, J.M., Cortés-Domínguez, I. and Ortiz-de-Solórzano, C. (2020) NMF-RI: Blind Spectral Unmixing of Highly Mixed Multispectral Flow and Image Cytometry Data. Bioinformatics, 36, 1590-1598. &gt;https://doi.org/10.1093/bioinformatics/btz751 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pengo, T., Muñoz-Barrutia, A., Zudaire, I. and Ortiz-de-Solorzano, C. (2013) Efficient Blind Spectral Unmixing of Fluorescently Labeled Samples Using Multi-Layer Non-Negative Matrix Factorization. PLOS ONE, 8, e78504. &gt;https://doi.org/10.1371/journal.pone.0078504 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Okada, S., Ohzeki, M. and Taguchi, S. (2019) Efficient Partition of Integer Optimization Problems with One-Hot Encoding. Scientific Reports, 9, Article No. 13036. &gt;https://doi.org/10.1038/s41598-019-49539-6 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhang, Y., Gong, L. and Wang, Y. (2005) An Improved TF-IDF Approach for Text Classification. Journal of Zhejiang University Science, 6, 49-55. &gt;https://doi.org/10.1631/jzus.2005.a49 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bafna, P., Pramod, D. and Vaidya, A. (2016) Document Clustering: TF-IDF Approach. 2016 International Conference on Electrical, Electronics, and Optimization Techniques (ICEEOT), Chennai, 3-5 March 2016, 61-66. &gt;https://doi.org/10.1109/iceeot.2016.7754750
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hien, L.T.K. and Gillis, N. (2021) Algorithms for Nonnegative Matrix Factorization with the Kullback-Leibler Divergence. Journal of Scientific Computing, 87, Article 93.
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Berry, M.W., Gillis, N. and Glineur, F. (2009) Document Classification Using Nonnegative Matrix Factorization and Underapproximation. 2009 IEEE International Symposium on Circuits and Systems, Taipei City, 24-27 May 2009, 2782-2785. &gt;https://doi.org/10.1109/iscas.2009.5118379 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref24">
    <label>24</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hu, L., Wu, N. and Li, X. (2022) Feature Nonlinear Transformation Non-Negative Matrix Factorization with Kullback-Leibler Divergence. Pattern Recognition, 132, Article 108906. &gt;https://doi.org/10.1016/j.patcog.2022.108906 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref25">
    <label>25</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nkurunziza, J.P., Nahayo, F. and Gillis, N. (2024) Orthogonal Nonnegative Matrix Factorization with the Kullback–Leibler Divergence. Pattern Recognition Letters, 197, 353-358. &gt;https://doi.org/10.1016/j.patrec.2025.08.012 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref26">
    <label>26</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hiemstra, D. (2000) A Probabilistic Justification for Using TF×IDF Term Weighting in Information Retrieval. International Journal on Digital Libraries, 3, 131-139.
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref27">
    <label>27</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Losada, D.E. and Azzopardi, L. (2008) An Analysis on Document Length Retrieval Trends in Language Modeling Smoothing. Information Retrieval, 11, 109-138. &gt;https://doi.org/10.1007/s10791-007-9040-x.
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref28">
    <label>28</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hazem, A. and Morin, E. (2013) A Comparison of Smoothing Techniques for Bilingual Lexicon Extraction from Comparable Corpora. Proceedings of the Sixth Workshop on Building and Using Comparable Corpora, Sofia, August 2013, 24-33.
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref29">
    <label>29</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Katz, S. (1987) Estimation of Probabilities from Sparse Data for the Language Model Component of a Speech Recognizer. IEEE Transactions on Acoustics, Speech, and Signal Processing, 35, 400-401. &gt;https://doi.org/10.1109/tassp.1987.1165125
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref30">
    <label>30</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jayaweera, M.P. and Dias, N.G.J. (2015) Application of Witten-Bell Discounting Techniques for Smoothing in Part of Speech Tagging Algorithm for Sinhala Language. Proceedings of the International Postgraduate Research Conference, Sri Lanka, 10 June 2015, p. 167. 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref31">
    <label>31</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Qaiser, S. and Ali, R. (2018) Text Mining: Use of TF-IDF to Examine the Relevance of Words to Documents. International Journal of Computer Applications, 181, 25-29. &gt;https://doi.org/10.5120/ijca2018917395
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref32">
    <label>32</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Manning, C.D., Raghavan, P. and Schütze, H. (2008) Term Weighting, and the Vector Space Model. In: Introduction to Information Retrieval, Cambridge University Press, 109-133.
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref33">
    <label>33</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kikuchi, M., Yoshida, M., Okabe, M. and Umemura, K. (2015) Confidence Interval of Probability Estimator of Laplace Smoothing. 2015 2nd International Conference on Advanced Informatics: Concepts, Theory and Applications (ICAICTA), Chonburi, 19-22 August 2015, 1-6. &gt;https://doi.org/10.1109/icaicta.2015.7335387
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref34">
    <label>34</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Van Laarhoven, T. (2017) L2 Regularization versus Batch and Weight Normalization. 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref35">
    <label>35</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gillis, N. (2020) Nonnegative Matrix Factorization. Society for Industrial and Applied Mathematics. &gt;https://doi.org/10.1137/1.9781611976410 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref36">
    <label>36</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Gillis, N. (2014) Successive Nonnegative Projection Algorithm for Robust Nonnegative Blind Source Separation. SIAM Journal on Imaging Sciences, 7, 1420-1450. &gt;https://doi.org/10.1137/130946782 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref37">
    <label>37</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Singer, S. and Nelder, J. (2009) Nelder-Mead Algorithm. Scholarpedia, 4, Article 2928. &gt;https://doi.org/10.4249/scholarpedia.2928 
    </mixed-citation>
   </ref>
   <ref id="scirp.146377-ref38">
    <label>38</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhong, S. and Ghosh, J. (2005) Generative Model-Based Document Clustering: A Comparative Study. Knowledge and Information Systems, 8, 374-384. &gt;https://doi.org/10.1007/s10115-004-0194-1
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>