<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ojapps
   </journal-id>
   <journal-title-group>
    <journal-title>
     Open Journal of Applied Sciences
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2165-3917
   </issn>
   <issn publication-format="print">
    2165-3925
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ojapps.2025.158153
   </article-id>
   <article-id pub-id-type="publisher-id">
    ojapps-144870
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Biomedical 
     </subject>
     <subject>
       Life Sciences, Chemistry 
     </subject>
     <subject>
       Materials Science, Computer Science 
     </subject>
     <subject>
       Communications, Engineering, Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    OTE-24LD: An Extended Descriptor Integrating Long-Distance Correlations for the Prediction of Macromolecular Interactions
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Obonan Etienne
      </surname>
      <given-names>
       Traore
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Ndiffon Charlemagne
      </surname>
      <given-names>
       Kopoin
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Dagou Dangui Augustin Sylvain Legrand
      </surname>
      <given-names>
       Koffi
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Gbame Gbede
      </surname>
      <given-names>
       Sylvain
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff3"> 
      <sup>3</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Souleymane
      </surname>
      <given-names>
       Oumtanaga
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aLaboratoire des Sciences de Données et Intelligence Artificielle (LASDIA), École Doctorale Polytechnique des Sciences et Technologies de l’Ingénieur (EDP-STI), Institut National Polytechnique Félix Houphouët-Boigny (INP-HB), Yamoussoukro, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aÉcole Supérieure Africaine des Technologies de l’Information et de la Communication (ESATIC), Abidjan, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <aff id="aff3">
    <addr-line>
     aUniversité Félix Houphouët-Boigny (UFHB), UFR Mathématiques-Informatique, Département Informatique, Abidjan, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     01
    </day> 
    <month>
     08
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    15
   </volume> 
   <issue>
    08
   </issue>
   <fpage>
    2308
   </fpage>
   <lpage>
    2318
   </lpage>
   <history>
    <date date-type="received">
     <day>
      25,
     </day>
     <month>
      July
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      15,
     </day>
     <month>
      July
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      15,
     </day>
     <month>
      August
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    The prediction of interactions between biological macromolecules, particularly macromolecular interactions, remains a major challenge in structural and functional bioinformatics. Numerous feature extraction methods have been developed, relying primarily on the physicochemical properties of amino acids and their sequential relationships to address this issue. Among these approaches, descriptors such as AAC (Amino Acid Composition), DPC (Dipeptide Composition), CTD (Composition-Transition-Distribution), and PseAAC (Pseudo Amino Acid Composition) have been widely used to transform macromolecular sequences into numerical vectors suitable for machine learning models. In this context, the OTE-24 method was recently introduced to capture local correlations between residues based on two normalized physicochemical properties. Despite its performance, as with several earlier methods, this model suffers from an intrinsic limitation: it overlooks long-distance correlations, which play a crucial role in the formation of recognition sites and the stability of macromolecular complexes. To address this limitation, we propose in this study an optimized extension of OTE-24, named OTE-24LD (Long Distance), which enhances the original descriptor by integrating distant relationships between residues, using a decreasing weighting factor modulated by positional distance. This improvement makes it possible to capture functional interactions often missed by descriptors based solely on immediate neighborhoods. The evaluation of this method, conducted on the HPRD dataset, demonstrates a 6.73% improvement in precision compared to earlier approaches.
   </abstract>
   <kwd-group> 
    <kwd>
     Macromolecular Interactions
    </kwd> 
    <kwd>
      Feature Extraction
    </kwd> 
    <kwd>
      Long-Distance Correlations
    </kwd> 
    <kwd>
      OTE-24LD Descriptor
    </kwd> 
    <kwd>
      Protein-Protein Interaction Prediction
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>The prediction of macromolecular interactions is a key component of structural bioinformatics, with major applications in systems biology, drug discovery, and protein engineering. To effectively model these interactions, it is crucial to represent protein sequences as numerical vectors that can be processed by machine learning algorithms <xref ref-type="bibr" rid="scirp.144870-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.144870-2">
     [2]
    </xref>. Descriptors based on the physicochemical properties of amino acids have thus been widely employed to capture relevant sequence features <xref ref-type="bibr" rid="scirp.144870-3">
     [3]
    </xref>. Among the classical descriptors are Amino Acid Composition (AAC), Dipeptide Composition (DPC), Composition-Transition-Distribution (CTD), and Pseudo Amino Acid Composition (PseAAC) <xref ref-type="bibr" rid="scirp.144870-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.144870-5">
     [5]
    </xref>. These methods primarily focus on the local or global properties of sequences but present limitations, particularly regarding the consideration of long-distance interactions between residues. Recent studies have highlighted the importance of integrating structural information and long-range correlations to improve the accuracy of protein function and interaction predictions <xref ref-type="bibr" rid="scirp.144870-6">
     [6]
    </xref> <xref ref-type="bibr" rid="scirp.144870-7">
     [7]
    </xref>.</p>
   <p>In this context, the OTE-24 method was proposed to enhance the representation of macromolecular sequences by integrating two normalized physicochemical properties and capturing local correlations between adjacent residues <xref ref-type="bibr" rid="scirp.144870-8">
     [8]
    </xref>. Although this approach has demonstrated improved performance compared to traditional descriptors, it remains limited by its focus on immediate local interactions, thereby neglecting long-distance correlations, which are often critical to protein function and structure <xref ref-type="bibr" rid="scirp.144870-9">
     [9]
    </xref>-<xref ref-type="bibr" rid="scirp.144870-11">
     [11]
    </xref>.</p>
   <p>To overcome this limitation, we introduce in this study an extension of the OTE-24 method, named OTE-24LD (Long Distance). This new approach aims to integrate long-distance interactions between residues using a decreasing weighting factor based on the positional distance within the sequence. By capturing both local and distant correlations, OTE-24LD provides a more comprehensive and informative representation of protein sequences, potentially enhancing the prediction of interactions and biological functions.</p>
   <p>We evaluated the performance of OTE-24LD on several benchmark datasets for macromolecular interaction prediction derived from the work of Vazquez et al. <xref ref-type="bibr" rid="scirp.144870-12">
     [12]
    </xref>. To assess the effectiveness of our model, we employed the Random Forest algorithm.</p>
  </sec><sec id="s2">
   <title>2. OTE-24LD Approach</title>
   <sec id="s2_1">
    <title>2.1. Overview of OTE-24</title>
    <p>The OTE-24 method is a macromolecular sequence characterization technique based on the integration of the physicochemical properties of amino acids using a bigram approach. It involves computing matrices that represent the distances between amino acids according to properties such as hydrophobicity and hydrophilicity, and then extracting feature vectors by combining these matrices through correlation and concatenation methods. This approach captures both the local order of amino acids and their immediate contextual relationships within the sequence <xref ref-type="bibr" rid="scirp.144870-8">
      [8]
     </xref>.</p>
    <p>The OTE-24 approach offers the advantage of generating normalized, fixed-dimensional feature vectors, facilitating their use in classification models. It effectively exploits local relationships between amino acids, which is crucial for identifying functional or structural motifs. Furthermore, the combination of multiple physicochemical properties enhances the richness of the extracted information, contributing to improved prediction accuracy for molecular interactions.</p>
    <p>However, a notable limitation of this method is its limited capacity to model long-distance interactions between amino acids, which play a critical role in the three-dimensional structure and function of proteins. This restriction to the analysis of local relationships can reduce the model’s precision when predicting complex structural behaviors or interactions that require a more global overview of the sequence.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Motivation and Principles of OTE-24LD</title>
    <p>In macromolecular structures, it is common for amino acids that are distant within the sequence to interact directly within the three-dimensional structure. These interactions play a decisive role in the stability and function of biological macromolecules. However, the OTE-24 method, limited to successive amino acid pairs, is unable to capture these long-distance relationships. To address this limitation, it is necessary to introduce a mechanism capable of quantifying the correlations between the physicochemical property values of amino acids separated by a given distance. The proposed extension, named OTE-24LD, consists of enhancing the classical OTE-24 descriptor by integrating weighted correlations at various distances. The idea is to preserve the local information provided by OTE-24 while adding components that represent interactions between amino acids separated by a defined number of positions within the sequence. Taking these weighted correlations into account improves the model’s ability to capture the complexity of molecular structures and interactions at multiple scales.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Mathematical Formulation of OTE-24LD</title>
    <p>Let a protein sequence of length 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        L 
      </mi> 
     </math> be composed of amino acids denoted as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          A 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          A 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          A 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          A 
        </mi> 
        <mn>
          4 
        </mn> 
       </msub> 
      </mrow> 
     </math>, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mo>
        ⋯ 
      </mo> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          A 
        </mi> 
        <mi>
          L 
        </mi> 
       </msub> 
      </mrow> 
     </math>. Each amino acid 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          A 
        </mi> 
        <mi>
          k 
        </mi> 
       </msub> 
      </mrow> 
     </math> is associated with a numerical value for each considered physicochemical property, denoted as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             k 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             p 
           </mi> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>, where 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math> represents the property index (in this study, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         P 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         2 
       </mn> 
      </mrow> 
     </math>). The long-distance correlation between amino acids separated by 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        d 
      </mi> 
     </math> positions is calculated as follows:</p>
    <p>
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          V 
        </mi> 
        <mi>
          d 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <munderover> 
        <mstyle mathsize="140%" displaystyle="true"> 
         <mo>
           ∑ 
         </mo> 
        </mstyle> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mrow> 
         <mi>
           L 
         </mi> 
         <mo>
           − 
         </mo> 
         <mi>
           d 
         </mi> 
        </mrow> 
       </munderover> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            ω 
          </mi> 
          <mi>
            d 
          </mi> 
         </msub> 
         <mo>
           ⋅ 
         </mo> 
         <munderover> 
          <mstyle mathsize="140%" displaystyle="true"> 
           <mo>
             ∑ 
           </mo> 
          </mstyle> 
          <mrow> 
           <mi>
             p 
           </mi> 
           <mo>
             = 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mi>
            P 
          </mi> 
         </munderover> 
         <mtext>
             
         </mtext> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mrow> 
           <msub> 
            <mi>
              A 
            </mi> 
            <mi>
              k 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mi>
             p 
           </mi> 
          </mrow> 
         </msub> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mrow> 
           <msub> 
            <mi>
              A 
            </mi> 
            <mrow> 
             <mi>
               k 
             </mi> 
             <mo>
               + 
             </mo> 
             <mi>
               d 
             </mi> 
            </mrow> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mi>
             p 
           </mi> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         with 
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mn>
         1 
       </mn> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         k 
       </mi> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         L 
       </mi> 
       <mo>
         − 
       </mo> 
       <mn>
         1 
       </mn> 
      </mrow> 
     </math>(1)</p>
    <p>where:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        d 
      </mi> 
     </math> represents the distance between two amino acids, varying from 1 to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          D 
        </mi> 
        <mrow> 
         <mi>
           max 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>,</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          ω 
        </mi> 
        <mi>
          d 
        </mi> 
       </msub> 
      </mrow> 
     </math> is the weight associated with the distance d, and</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        P 
      </mi> 
     </math> is the number of physicochemical properties considered.</p>
    <p>This formulation makes it possible to quantify the interaction between all pairs of amino acids separated by 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        d 
      </mi> 
     </math> positions for each of the selected properties.</p>
    <p>To modulate the importance of interactions based on distance, we introduce a weighting factor 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          ω 
        </mi> 
        <mi>
          d 
        </mi> 
       </msub> 
      </mrow> 
     </math>, which follows an exponential decay defined as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          ω 
        </mi> 
        <mi>
          d 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msup> 
        <mtext>
          e 
        </mtext> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           α 
         </mi> 
         <mi>
           d 
         </mi> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>(2)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        α 
      </mi> 
     </math> is an adjustable parameter that controls the rate of decrease in the influence of interactions as the distance increases. This mechanism preserves the effect of nearby correlations while progressively reducing the influence of more distant relationships.</p>
    <p>For each macromolecular sequence, the 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          V 
        </mi> 
        <mi>
          d 
        </mi> 
       </msub> 
      </mrow> 
     </math> values are computed for all distances 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        d 
      </mi> 
     </math> ranging from 1 to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          D 
        </mi> 
        <mrow> 
         <mi>
           max 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>. The results are then concatenated to form a unique feature vector:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         V 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mn>
            3 
          </mn> 
         </msub> 
         <mo>
           ⋯ 
         </mo> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mrow> 
           <msub> 
            <mi>
              D 
            </mi> 
            <mrow> 
             <mi>
               max 
             </mi> 
            </mrow> 
           </msub> 
          </mrow> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mrow> 
           <mn>
             2 
           </mn> 
           <mo>
             , 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mrow> 
           <mn>
             2 
           </mn> 
           <mo>
             , 
           </mo> 
           <mn>
             2 
           </mn> 
          </mrow> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mrow> 
           <mn>
             2 
           </mn> 
           <mo>
             , 
           </mo> 
           <mn>
             3 
           </mn> 
          </mrow> 
         </msub> 
         <mo>
           ⋯ 
         </mo> 
         <msub> 
          <mi>
            V 
          </mi> 
          <mrow> 
           <mn>
             2 
           </mn> 
           <mo>
             , 
           </mo> 
           <mn>
             20 
           </mn> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math>(3)</p>
    <p>This vector thus brings together all the weighted long-distance correlations between amino acids for each of the selected properties. In this study, we considered two properties: hydrophobicity and hydrophilicity.</p>
   </sec>
   <sec id="s2_4">
    <title>2.4. Selection and Normalization of Physicochemical Properties</title>
    <p>For this OTE-24LD extension, two essential properties were selected. Hydrophobicity, a key property in the structuring of macromolecules and the formation of internal hydrophobic regions, and hydrophilicity, which influences the exposure of amino acids to solvents and interactions with the aqueous environment. These two properties were chosen because they are widely used in studies of macromolecular structural behavior.</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Experimental Protocol</title>
   <sec id="s3_1">
    <title>3.1. Dataset</title>
    <p>In this section, a class of possible solutions is proposed and the insertion of these solutions in the main model is investigated to check where they will lead to. This is initiated by the statement of the following claim.</p>
    <p>The study was conducted using data from the Human Protein Reference Database (HPRD), a reference resource for the functional and structural analysis of human proteins <xref ref-type="bibr" rid="scirp.144870-13">
      [13]
     </xref>. Developed through a collaboration between the Bioinformatics Institute of Bangalore (India) and the Pandey Laboratory at Johns Hopkins University (USA), HPRD compiles manually curated scientific annotations covering the majority of human macromolecules.</p>
    <p>The latest version of this dataset includes over 36,500 unique macromolecular interactions involving 25,000 proteins and 6360 isoforms <xref ref-type="bibr" rid="scirp.144870-14">
      [14]
     </xref> <xref ref-type="bibr" rid="scirp.144870-15">
      [15]
     </xref>.</p>
    <p>In this study, we maintained the same data proportions as the original OTE-24 method <xref ref-type="bibr" rid="scirp.144870-8">
      [8]
     </xref>. We performed preprocessing on the macromolecular sequences prior to feature extraction. First, incomplete and redundant sequences were filtered out to ensure the integrity and independence of the samples. Then, verification and homogenization of the macromolecular or protein alphabet al. allowed us to retain only sequences containing the 20 standard amino acids. Finally, the physicochemical properties of amino acids used in the descriptor calculations were normalized using the following transformation:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          C 
        </mi> 
        <mrow> 
         <mi>
           A 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           p 
         </mi> 
        </mrow> 
        <mrow> 
         <mtext>
           norm 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mrow> 
           <mi>
             A 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             p 
           </mi> 
          </mrow> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msubsup> 
          <mi>
            C 
          </mi> 
          <mi>
            p 
          </mi> 
          <mrow> 
           <mtext>
             min 
           </mtext> 
          </mrow> 
         </msubsup> 
        </mrow> 
        <mrow> 
         <msubsup> 
          <mi>
            C 
          </mi> 
          <mi>
            p 
          </mi> 
          <mrow> 
           <mtext>
             max 
           </mtext> 
          </mrow> 
         </msubsup> 
         <mo>
           − 
         </mo> 
         <msubsup> 
          <mi>
            C 
          </mi> 
          <mi>
            p 
          </mi> 
          <mrow> 
           <mtext>
             min 
           </mtext> 
          </mrow> 
         </msubsup> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(4)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mrow> 
         <mi>
           A 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           p 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is the raw value of a given property for amino acid A, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          C 
        </mi> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           min 
         </mi> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          C 
        </mi> 
        <mi>
          p 
        </mi> 
        <mrow> 
         <mi>
           max 
         </mi> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> are respectively the minimum and maximum values of this property across the 20 standard amino acids.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Experimental Parameters</title>
    <p>The parameters involved in the computation of long-range correlation descriptors were optimized based on preliminary experiments conducted on the HPRD dataset. A grid search procedure was employed to identify the most relevant combinations of values for the parameters 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          D 
        </mi> 
        <mrow> 
         <mi>
           max 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        α 
      </mi> 
     </math>, using a validation set extracted from the training data (80% of the HPRD dataset).</p>
    <p>The parameter 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          D 
        </mi> 
        <mrow> 
         <mi>
           max 
         </mi> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>, which defines the maximum distance between two amino acids considered for computing correlations between physicochemical properties, was evaluated over the range 10 to 30 with a step size of 5. The highest average performance was achieved for 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          D 
        </mi> 
        <mrow> 
         <mi>
           max 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mn>
         20 
       </mn> 
      </mrow> 
     </math>, which was therefore selected for the final model.</p>
    <p>Similarly, the exponential weighting factor 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        α 
      </mi> 
     </math>, which controls the attenuation of interaction influence as a function of distance, was explored within the interval [0.01, 0.10], with an increment of 0.01. The optimal value identified was 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         α 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         0.05 
       </mn> 
      </mrow> 
     </math>.</p>
    <p>With respect to the physicochemical properties considered, two fundamental attributes were retained: hydrophobicity and hydrophilicity, due to their crucial role in protein−protein interaction mechanisms.</p>
    <p>To ensure an unbiased performance assessment and minimize training-related biases, we adopted a stratified 5-fold cross-validation procedure in conjunction with the Random Forest algorithm. The HPRD dataset was randomly partitioned into five balanced subsets based on class labels (positive/negative). In each iteration, four subsets were used for training and the remaining one for testing. This process was repeated five times, with each subset serving once as the test set. Final performance metrics were computed as the average of the scores obtained across the five iterations.</p>
   </sec>
   <sec id="s3_3">
    <title>3.3. Comparison Methods</title>
    <p>To highlight the performance of the OTE-24LD method, it was compared to several classical descriptors from the literature, including AAC, DPC, CTD, PseAAC, and OTE-24. Recall that the AAC (Amino Acid Composition) model encodes the proportion of each amino acid within the sequence without order information <xref ref-type="bibr" rid="scirp.144870-16">
      [16]
     </xref> <xref ref-type="bibr" rid="scirp.144870-17">
      [17]
     </xref>. The DPC (Dipeptide Composition) model quantifies the frequency of successive amino acid pairs, thereby incorporating local order information <xref ref-type="bibr" rid="scirp.144870-18">
      [18]
     </xref> <xref ref-type="bibr" rid="scirp.144870-19">
      [19]
     </xref>. The CTD (Composition, Transition, Distribution) encodes sequences based on groupings of physicochemical properties and their distribution <xref ref-type="bibr" rid="scirp.144870-20">
      [20]
     </xref> <xref ref-type="bibr" rid="scirp.144870-21">
      [21]
     </xref>. The PseAAC (Pseudo Amino Acid Composition) enriches the AAC encoding by considering correlations between distant positions <xref ref-type="bibr" rid="scirp.144870-22">
      [22]
     </xref> <xref ref-type="bibr" rid="scirp.144870-23">
      [23]
     </xref>. Finally, the OTE-24 technique is a coding method based on weighted local interactions, considered here as the immediate reference method.</p>
    <p>We applied these different methods on the same HPRD dataset and used the same cross-validation procedure, thus ensuring a fair and rigorous comparison of their performance.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. Results and Discussion</title>
   <sec id="s4_1">
    <title>4.1. Evaluation Metrics for Overall Model Performance</title>
    <p>In this section, we present the performance of the proposed OTE-24LD model, trained using the Random Forest algorithm, for the prediction of interactions between biological macromolecules. The performance was assessed using several metrics relevant to the study context. These metrics reflect different aspects of the model’s predictive behavior, particularly in the context of imbalanced classes. We used accuracy to provide a general measure of prediction correctness, while precision, recall, and specificity assess the model’s ability to correctly identify positive and negative interactions. The F1-score offers a meaningful balance between precision and recall, especially useful in imbalanced class scenarios. Balanced accuracy corrects for the impact of class imbalance, and the Matthews Correlation Coefficient (MCC) provides a robust evaluation even when classes are unequal. Additionally, the ROC and PR curves, along with their respective areas under the curve (AUC and AUPR), enable graphical and quantitative analysis of the model’s performance across various classification thresholds. The confusion matrix completes this analysis by offering a detailed visualization of prediction errors. These metrics were selected to ensure a rigorous and multidimensional evaluation of performance, tailored to the specific contextual characteristics of the problem studied. They are calculated as follows:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Accuracy 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           TN 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           TN 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FN 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(5)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Precision 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FP 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(6)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Recall 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FN 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(7)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         specificity 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TN 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TN 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FP 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(8)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         F 
       </mtext> 
       <mn>
         1 
       </mn> 
       <mtext>
         -score 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mn>
         2 
       </mn> 
       <mo>
         × 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Precision 
         </mtext> 
         <mo>
           × 
         </mo> 
         <mtext>
           Sensibilite 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           Precision 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           Sensibilite 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(9)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Balanced Accuracy 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Sensibilite 
         </mtext> 
         <mo>
           × 
         </mo> 
         <mtext>
           Specificite 
         </mtext> 
        </mrow> 
        <mn>
          2 
        </mn> 
       </mfrac> 
      </mrow> 
     </math>(10)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         MCC 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             TP 
           </mtext> 
           <mo>
             × 
           </mo> 
           <mtext>
             TN 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           − 
         </mo> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             FP 
           </mtext> 
           <mo>
             × 
           </mo> 
           <mtext>
             FN 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <msqrt> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mtext>
               TP 
             </mtext> 
             <mo>
               + 
             </mo> 
             <mtext>
               FP 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mtext>
               TP 
             </mtext> 
             <mo>
               + 
             </mo> 
             <mtext>
               FN 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mtext>
               TN 
             </mtext> 
             <mo>
               + 
             </mo> 
             <mtext>
               FP 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mtext>
               TN 
             </mtext> 
             <mo>
               + 
             </mo> 
             <mtext>
               FN 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msqrt> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math>(11)</p>
    <p>Together, these indicators allow for a comprehensive analysis of the model’s performance, both in terms of overall prediction accuracy and the differentiation between positive and negative classes. The following table summarizes the performances obtained across all test sets (<xref ref-type="table" rid="table1">
      Table 1
     </xref>).</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.144870-"></xref>Table 1. Overall performance of the OTE-24LD model using random forest.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Metrics</p></td> 
       <td class="custom-bottom-td acenter" width="22.61%"><p style="text-align:center">Values (%)</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">Accuracy</p></td> 
       <td class="custom-top-td acenter" width="22.61%"><p style="text-align:center">93.51</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">Precision</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">96.73</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">Recall</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">90.10</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">Specificity</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">96.94</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">F1-Score</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">93.30</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">Balanced Accuracy</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">93.52</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">Matthews Correlation Coefficient</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">87.23</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">AUC (ROC)</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">97.99</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="22.61%"><p style="text-align:center">AUPR</p></td> 
       <td class="acenter" width="22.61%"><p style="text-align:center">98.17</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The model achieved an accuracy of 93.51%, indicating that a large proportion of the total predictions were correct. However, in contexts where class imbalance exists, accuracy alone can be misleading. For this reason, balanced accuracy was also computed, reaching 93.52%, reflecting a balanced ability to correctly predict both classes. The recall of 90.10% demonstrates the model’s ability to correctly identify positive cases, while the specificity of 96.94% highlights its capacity to avoid false positives. The F1-score, which synthesizes precision and recall, reached 93.30%, indicating a good trade-off between interaction detection and error minimization. The Matthews Correlation Coefficient (MCC), recognized for its robustness in imbalanced class scenarios, achieved 87.23%, further confirming the overall quality of the model.</p>
   </sec>
   <sec id="s4_2">
    <title>4.2. Comparison with Existing Methods</title>
    <p>To assess the effectiveness of our proposed approach, we conducted a comparative evaluation against several widely recognized models from the literature, including AAC, DPC, and PseAAC, each coupled with different classifiers. The original OTE-24 model, primarily designed to capture short-range Bigram interactions, was used as a baseline in our experiments. It achieved an accuracy of 86.61%, an F1-score of 87.33%, a ROC-AUC of 89.43%, and a Matthews Correlation Coefficient (MCC) of 86.19. In contrast, our enhanced method, OTE-24LD, which incorporates both physicochemical properties and long-range interaction features, demonstrates a 6.73% improvement in accuracy over OTE-24.</p>
    <p>This performance gain is further supported by the results presented in the table below, which compares our approach to conventional methods that also consider long-distance interactions (<xref ref-type="table" rid="table2">
      Table 2
     </xref>).</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.144870-"></xref>Table 2. Performance comparison between OTE-24LD and other approaches.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="16.69%"><p style="text-align:center">Approach</p></td> 
       <td class="custom-bottom-td acenter" width="16.67%"><p style="text-align:center">Classifier</p></td> 
       <td class="custom-bottom-td acenter" width="16.66%"><p style="text-align:center">Accuracy (%)</p></td> 
       <td class="custom-bottom-td acenter" width="16.66%"><p style="text-align:center">F1-Score (%)</p></td> 
       <td class="custom-bottom-td acenter" width="18.38%"><p style="text-align:center">ROC-AUC (%)</p></td> 
       <td class="custom-bottom-td acenter" width="14.94%"><p style="text-align:center">MCC</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="16.69%"><p style="text-align:center">AAC</p></td> 
       <td class="custom-top-td acenter" width="16.67%"><p style="text-align:center">SVM</p></td> 
       <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">69.20</p></td> 
       <td class="custom-top-td acenter" width="16.66%"><p style="text-align:center">80.60</p></td> 
       <td class="custom-top-td acenter" width="18.38%"><p style="text-align:center">58.00</p></td> 
       <td class="custom-top-td acenter" width="14.94%"><p style="text-align:center">0.10</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.69%"><p style="text-align:center">DPC</p></td> 
       <td class="acenter" width="16.67%"><p style="text-align:center">RF</p></td> 
       <td class="acenter" width="16.66%"><p style="text-align:center">70.70</p></td> 
       <td class="acenter" width="16.66%"><p style="text-align:center">82.50</p></td> 
       <td class="acenter" width="18.38%"><p style="text-align:center">61.40</p></td> 
       <td class="acenter" width="14.94%"><p style="text-align:center">0.06</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.69%"><p style="text-align:center">PseAAC</p></td> 
       <td class="acenter" width="16.67%"><p style="text-align:center">SVM</p></td> 
       <td class="acenter" width="16.66%"><p style="text-align:center">65.40</p></td> 
       <td class="acenter" width="16.66%"><p style="text-align:center">76.80</p></td> 
       <td class="acenter" width="18.38%"><p style="text-align:center">58.20</p></td> 
       <td class="acenter" width="14.94%"><p style="text-align:center">0.09</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="16.69%"><p style="text-align:center">OTE-24LD</p></td> 
       <td class="acenter" width="16.67%"><p style="text-align:center">RF</p></td> 
       <td class="acenter" width="16.66%"><p style="text-align:center">93.57</p></td> 
       <td class="acenter" width="16.66%"><p style="text-align:center">93.30</p></td> 
       <td class="acenter" width="18.38%"><p style="text-align:center">97.99</p></td> 
       <td class="acenter" width="14.94%"><p style="text-align:center">0.8723</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The OTE-24LD method outperforms classical approaches in terms of accuracy, F1-score, and AUC, confirming the value of integrating physicochemical properties and long-distance interactions.</p>
   </sec>
   <sec id="s4_3">
    <title>4.3. Discussion</title>
    <p>The OTE-24LD (Long Distance) method stands out from existing approaches in the literature by explicitly integrating long-range correlations between residues. While conventional descriptors such as AAC and PseAAC limited to local frequencies or immediate neighborhoods struggle to exceed 70% accuracy, OTE-24LD achieves a global accuracy of 93.51%. This significant improvement highlights the importance of capturing extended sequential dependencies in the prediction of macromolecular interactions.</p>
    <p>This enhancement is also reflected in the F1-score, which increases from 82.50% for DPC (used as the reference classifier) to 93.30% with OTE-24LD. Such a gain demonstrates a better balance between recall and precision, made possible by our adaptive weighting scheme. This mechanism preserves the influence of distant residues without introducing redundancy or diluting relevant local information. The AUC reaches 97.99%, a value close to the theoretical optimum, indicating an excellent trade-off between sensitivity and specificity. Likewise, the Matthews Correlation Coefficient (MCC), a metric on which traditional methods often underperform in this context, reaches 0.8723, reflecting a well-calibrated and reliable model capable of limiting both false positives and false negatives.</p>
    <p>Beyond these quantitative results, the ability to capture spatially distant but functionally coupled residues enables more accurate prediction of interaction interfaces, better modeling of complex stability, and identification of key regions involved in molecular recognition. This level of precision contributes to a deeper understanding of interaction mechanisms and supports the rational design of therapeutic biomolecules by reducing dependence on costly and time-consuming in vitro experiments.</p>
    <p>Nevertheless, the OTE-24LD approach presents certain limitations that should be acknowledged. It relies solely on information derived from the primary sequence and does not incorporate structural or evolutionary features, which could further refine predictive performance. Additionally, computing long-distance correlation descriptors incurs a non-negligible computational cost. The simulations conducted in this study were performed on a high-performance server (25.1 GHz CPU, 128 GB RAM, 830 GB SSD storage). Although resource consumption remained within reasonable limits, this requirement may constitute a constraint for large-scale or real-time applications without further optimization.</p>
    <p>By reducing reliance on experimental validation, OTE-24LD has the potential to accelerate the development and refinement of therapeutic candidates. This approach offers a robust, relevant, and generalizable solution for leveraging long-range sequence information in the prediction of macromolecular interactions. Through its performance, interpretability, and practical scope, it represents a promising advancement for research in structural and functional bioinformatics.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Conclusions and Suggestions</title>
   <p>This study enabled us to develop a predictive model for macromolecular interactions based on the OTE-24LD vector, enhanced by the integration of long-distance correlations and normalized physicochemical properties of amino acids. The simulations conducted on the HPRD database confirm the superiority of our approach, notably achieving an accuracy of 93.51%, an F1-score of 93.30%, an AUC of 97.99%, and an MCC of 0.8723. These metric values reflect the model’s reinforced ability to accurately distinguish true interactions from false signals, well beyond the performance levels observed with traditional methods such as AAC, DPC, or PseAAC.</p>
   <p>The introduction of a decreasing weighting factor for distant residues within the sequence allowed the model to capture long-range functional patterns, often underestimated by conventional descriptors. This strategy highlights the critical importance of considering distance effects in the prediction of macromolecular interfaces, offering a comparative precision gain of over 20% and achieving near-ideal discrimination according to ROC metric values.</p>
   <p>Beyond validation on the HPRD dataset, this work lays the foundation for promising future extensions. On the one hand, incorporating 3D structural data could further enrich the OTE-24LD method; on the other hand, applying deep learning techniques or hybrid ensemble approaches to the generated feature vectors opens the way for even more robust predictive models.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.144870-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Char, S., Corley, N., Alamdari, S., Yang, K.K. and Amini, A.P. (2025) ProtNote: A Multimodal Method for Protein-Function Annotation. Bioinformatics, 41, btaf170. &gt;https://doi.org/10.1093/bioinformatics/btaf170
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Taha, K. (2025) Protein-Protein Interaction Detection Using Deep Learning: A Survey, Comparative Analysis, and Experimental Evaluation. Computers in Biology and Medicine, 185, Article ID: 109449. &gt;https://doi.org/10.1016/j.compbiomed.2024.109449
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kiouri, D.P., Batsis, G.C. and Chasapis, C.T. (2025) Structure-Based Approaches for Protein-Protein Interaction Prediction Using Machine Learning and Deep Learning. Biomolecules, 15, Article 141. &gt;https://doi.org/10.3390/biom15010141
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Shen, H. and Chou, K. (2008) Pseaac: A Flexible Web Server for Generating Various Kinds of Protein Pseudo Amino Acid Composition. Analytical Biochemistry, 373, 386-388. &gt;https://doi.org/10.1016/j.ab.2007.10.012
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chou, K. and Cai, Y. (2003) Predicting Protein Quaternary Structure by Pseudo Amino Acid Composition. Proteins: Structure, Function, and Bioinformatics, 53, 282-289. &gt;https://doi.org/10.1002/prot.10500
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Emonts, J. and Buyel, J.F. (2023) An Overview of Descriptors to Capture Protein Properties—Tools and Perspectives in the Context of QSAR Modeling. Computational and Structural Biotechnology Journal, 21, 3234-3247. &gt;https://pmc.ncbi.nlm.nih.gov/articles/PMC10781719/?utm_source=chatgpt.com
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Tang, T., Li, T., Li, W., Cao, X., Liu, Y. and Zeng, X. (2024) Anti-Symmetric Framework for Balanced Learning of Protein-Protein Interactions. Bioinformatics, 40, btae603. &gt;https://doi.org/10.1093/bioinformatics/btae603
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Obonan Etienne, T., Ndiffon Charlemagne, K., Tchimou Guepie Euloge, N. and Souleymane, O. (2025) Optimization of Feature Extraction for the Prediction of Macromolecular Interactions: OTE-24 Approach. International Journal of Advanced Research, 13, 577-589. &gt;https://doi.org/10.21474/ijar01/20598
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Hosseini, S., Golding, G.B. and Ilie, L. (2024) Seq-Insite: Sequence Supersedes Structure for Protein Interaction Site Prediction. Bioinformatics, 40, btad738. &gt;https://doi.org/10.1093/bioinformatics/btad738
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Detlefsen, N.S., Hauberg, S. and Boomsma, W. (2022) Learning Meaningful Representations of Protein Sequences. Nature Communications, 13, Article No. 1914. &gt;https://doi.org/10.1038/s41467-022-29443-w
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cao, M., Zainudin, S. and Daud, K.M. (2025) Feature Fusion with Attributed Deepwalk for Protein-Protein Interaction Prediction. Scientific Reports, 15, Article No. 12255. &gt;https://doi.org/10.1038/s41598-025-96510-9
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, G., Liu, X., Wang, K., Gao, Y., Li, G., Baptista-Hon, D.T., et al. (2023) Deep-Learning-Enabled Protein-Protein Interaction Analysis for Prediction of SARS-CoV-2 Infectivity and Variant Evolution. Nature Medicine, 29, 2007-2018. &gt;https://doi.org/10.1038/s41591-023-02483-5
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kandasamy, R.K., et al. (2025) Human Proteinpedia: A Unified Discovery Resource for Proteomics Research. Nucleic Acids Research, 37, D773-D781.&gt;https://www.researchgate.net/publication/23410255_Human_Proteinpedia_A_unified_discovery_resource_for_proteomics_research 
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Goel, R., Harsha, H.C., Pandey, A. and Prasad, T.S.K. (2012) Human Protein Reference Database and Human Proteinpedia as Resources for Phosphoproteome Analysis. Molecular BioSystems, 8, 453-463. &gt;https://doi.org/10.1039/c1mb05340j
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Keshava Prasad, T.S., Goel, R., Kandasamy, K., Keerthikumar, S., Kumar, S., Mathivanan, S., et al. (2009) Human Protein Reference Database—2009 Update. Nucleic Acids Research, 37, D767-D772. &gt;https://doi.org/10.1093/nar/gkn892
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Breimann, S. and Frishman, D. (2024) AAclust: k-Optimized Clustering for Selecting Redundancy-Reduced Sets of Amino Acid Scales. Bioinformatics Advances, 4, vbae165. &gt;https://doi.org/10.1093/bioadv/vbae165
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jiao, S., Ye, X., Sakurai, T., Zou, Q. and Liu, R. (2024) Integrated Convolution and Self-Attention for Improving Peptide Toxicity Prediction. Bioinformatics, 40, btae297. &gt;https://doi.org/10.1093/bioinformatics/btae297
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ullah, F., Salam, A., Nadeem, M., Amin, F., AlSalman, H., Abrar, M., et al. (2024) Extended Dipeptide Composition Framework for Accurate Identification of Anticancer Peptides. Scientific Reports, 14, Article No. 17381. &gt;https://doi.org/10.1038/s41598-024-68475-8
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yan, C., Geng, A., Pan, Z., Zhang, Z. and Cui, F. (2024) MultiFeatVotPIP: A Voting-Based Ensemble Learning Framework for Predicting Proinflammatory Peptides. Briefings in Bioinformatics, 25, bbae505. &gt;https://doi.org/10.1093/bib/bbae505
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ghafoor, H., Abbasi, A.F., Asim, M.N. and Dengel, A. (2024) CTD-Global (CTD-G): A Novel Composition, Transition, and Distribution Based Peptide Sequence Encoder for Hormone Peptide Prediction. Informatics in Medicine Unlocked, 50, Article ID: 101578. &gt;https://doi.org/10.1016/j.imu.2024.101578
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Akbar, S., Raza, A. and Zou, Q. (2024) Deepstacked-AVPs: Predicting Antiviral Peptides Using Tri-Segment Evolutionary Profile and Word Embedding Based Multi-Perspective Features with Deep Stacking Model. BMC Bioinformatics, 25, Article No. 102. &gt;https://doi.org/10.1186/s12859-024-05726-5
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref22">
    <label>22</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Qiu, W., Xiao, X., Lin, W. and Chou, K. (2014) iMethyl-PseAAC: Identification of Protein Methylation Sites via a Pseudo Amino Acid Composition Approach. BioMed Research International, 2014, Article ID: 947416. &gt;https://doi.org/10.1155/2014/947416
    </mixed-citation>
   </ref>
   <ref id="scirp.144870-ref23">
    <label>23</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chou, K. (2011) Some Remarks on Protein Attribute Prediction and Pseudo Amino Acid Composition. Journal of Theoretical Biology, 273, 236-247. &gt;https://doi.org/10.1016/j.jtbi.2010.12.024
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>