<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    cc
   </journal-id>
   <journal-title-group>
    <journal-title>
     Computational Chemistry
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2332-5968
   </issn>
   <issn publication-format="print">
    2332-5984
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/cc.2025.133003
   </article-id>
   <article-id pub-id-type="publisher-id">
    cc-143538
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Chemistry 
     </subject>
     <subject>
       Materials Science
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    QSAR Models: Exploring Limits in Three Cases
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       El Hadji Sawaliho
      </surname>
      <given-names>
       Bamba
      </given-names>
     </name>
    </contrib>
   </contrib-group> 
   <aff id="affnull">
    <addr-line>
     aConstitution and Reaction of Matter Laboratory, Training and Research Unit in Structural, Material and Technological Sciences, Felix Houphouet-Boigny University, Abidjan, Côte d’Ivoire
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     25
    </day> 
    <month>
     06
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    13
   </volume> 
   <issue>
    03
   </issue>
   <fpage>
    45
   </fpage>
   <lpage>
    68
   </lpage>
   <history>
    <date date-type="received">
     <day>
      9,
     </day>
     <month>
      May
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      22,
     </day>
     <month>
      May
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      22,
     </day>
     <month>
      June
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    This article critically assessed the validity of five multiple linear regression models across three separate studies. The first examined the cytotoxic properties of N-tosyl-1,2,3,4-tetrahydroisoquinoline compounds. The second evaluated the antiproliferative effects of 1,3,5-arylidene rhodanines. The last explored the antitumour potential of thiazoline or thiazine derivatives. Despite limited sample sizes, the model validation showed robust performance and predictive capabilities. However, their forecasts lacked accuracy. The authors validated their models by assessing the fit training data and generalization ability. The gaps weren’t clearly defined, and outliers were only partially considered. The cytotoxicity study of N-Tosyl-1,2,3,4-Tetrahydroisoquinoline used a ±2 standardized residual. Non-random sampling can introduce selection bias. Ignoring dispersion and employing fixed molecules can reduce model accuracy. Adhering to MLR premises aids in validation. Analysis of secondary data from three articles showed that all five MLR models were invalid, emphasizing the need to verify MLR assumptions before utilizing the QSAR approach.
   </abstract>
   <kwd-group> 
    <kwd>
     Multiple Linear Regression
    </kwd> 
    <kwd>
      QSAR
    </kwd> 
    <kwd>
      Tetrahydroisoquinoline
    </kwd> 
    <kwd>
      Arylidene Rhodanine
    </kwd> 
    <kwd>
      Thiazoline
    </kwd> 
    <kwd>
      Thiazine
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>Research in chemistry, pharmaceuticals, and therapeutics renewed interest in Quantitative Structure-Activity Relationship (QSAR) modelling, which predicts biological activity of molecules based on their structure. Multiple Linear Regression (MLR) is used to link these two parts in this approach. The Dependent Variable (DV) is the amount needed for a response, while Independent Variables (IV) are the compounds’ properties. QSAR’s assumption was that a similar configuration has analogous activities, making it a powerful tool across scientific fields. Models were employed in drug discovery to project the effectiveness and safety of new molecules. Their performances were assessed using the coefficient of determination (R<sup>2</sup>), which indicated how well they fit the training data. Prediction accuracy was measured by Root Mean Square Error of the Training (RMSTr), set and by Root Means Square Error of Testing set. The correlation coefficient of Cross-Validation 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msubsup> 
       <mi>
         Q 
       </mi> 
       <mrow> 
        <mi>
          C 
        </mi> 
        <mi>
          V 
        </mi> 
       </mrow> 
       <mn>
         2 
       </mn> 
      </msubsup> 
     </mrow> 
    </math> evaluated the predictive power of Model. This work selected three articles to assess the relevance of the standards used in several QSAR-based research studies. Furthermore, these papers proposed models aimed at projecting the activities of molecular families against serious diseases such as cancer. They offered the opportunity to judge the conformity of their models to the premises underlying a valid MLR. One key factor behind this decision was the small sample size, which may have resulted in the breach of certain MLR assumptions <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref>. The standards affecting the corroboration of the three studies’ models were investigated.</p>
   <p>Case 1 analyzed the cytotoxicity of 12 N-Tosyl-1,2,3,4-Tetrahydroisoquinoline molecules <xref ref-type="bibr" rid="scirp.143538-2">
     [2]
    </xref>. The models explained 79.19% of the variance for DV MOLT3 (Model 1) and 98.49% for DV HepG2 (Model 2). Model 1 demonstrated a predictive power of 56.9%, with forecasts’ accuracy of 22.3% in training and 32.5% in testing. Model 2 predicted with 91.5% power. Its precision was 2.8% in training and 7.3% in testing. Model 1 exhibited high performance with low prediction accuracy and reliability. Model 2 showed excellent performance and prediction power, but its precision remains low. These statistics suggest that a model’s performance and projection capabilities (precision and power) can be contradictory. They noted that using a high R<sup>2</sup> to validate a model has limitations. The small sample sizes (N = 12 for DV MOLT3, N = 7 for DV HepG2) may reduce models’ predictive power <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-3">
     [3]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref>. Model 1’s low predictive capacity is evident. Model 2’s high value is questionable due to its smaller sample size (N = 7 vs. N = 12).</p>
   <p>Case 2 evaluated the antiproliferative activity of 13 1,3,5-Arylidene Rhodanine compounds <xref ref-type="bibr" rid="scirp.143538-5">
     [5]
    </xref>. The performance was 92.7% for Model 3 (ActMDA) and 88.2% for Model 4 (ActNCI). The predictive power of Model 3 was 95.4%, while that of Model 4 was 92.6%, demonstrating their overall effectiveness. However, their sample sizes (N = 13) hardly explained this level of the two models’ external validity. The absence of precise forecast data hindered the evaluation of this key component in both models.</p>
   <p>Case 3 aimed to predict pulmonary effects of Thiazine or Thiazoline on A-549 cells <xref ref-type="bibr" rid="scirp.143538-6">
     [6]
    </xref>. Model 5’s performance reached 90.5%, with a prediction accuracy of 10.6% for the test sample. Its predictive power was also 90.5%. It was consistent with the performance. Nevertheless, the low reliability of its predictions didn’t support it. Furthermore, the small sample size (N = 14) is difficult to reconcile with predictive power or high generalizability of the results <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-3">
     [3]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref>.</p>
   <p>The analyzed cases showed that their models’ performances were elevated, effectively accounting for the observed variances. Their predictive powers were high (except for Model 1). Conversely, they weren’t precise. Reference <xref ref-type="bibr" rid="scirp.143538-2">
     [2]
    </xref> <xref ref-type="bibr" rid="scirp.143538-5">
     [5]
    </xref> <xref ref-type="bibr" rid="scirp.143538-6">
     [6]
    </xref> employed the criterion 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msup> 
       <mi>
         R 
       </mi> 
       <mn>
         2 
       </mn> 
      </msup> 
      <mo>
        − 
      </mo> 
      <msubsup> 
       <mi>
         Q 
       </mi> 
       <mrow> 
        <mi>
          C 
        </mi> 
        <mi>
          V 
        </mi> 
       </mrow> 
       <mn>
         2 
       </mn> 
      </msubsup> 
      <mo>
        &lt; 
      </mo> 
      <mn>
        0.3 
      </mn> 
     </mrow> 
    </math> to validate their models. This criterion included two parameters: 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msup> 
       <mi>
         R 
       </mi> 
       <mn>
         2 
       </mn> 
      </msup> 
     </mrow> 
    </math> evaluating their fit to the training data and 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msubsup> 
       <mi>
         Q 
       </mi> 
       <mrow> 
        <mi>
          C 
        </mi> 
        <mi>
          V 
        </mi> 
       </mrow> 
       <mn>
         2 
       </mn> 
      </msubsup> 
     </mrow> 
    </math> assessing their ability to generalize to new molecules. The interpretation of the gap was unclear. Another weakness was the partial treatment of outliers <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref>.</p>
   <p>Reference <xref ref-type="bibr" rid="scirp.143538-7">
     [7]
    </xref> noted that Pearson’s 
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <msup> 
       <mi>
         R 
       </mi> 
       <mn>
         2 
       </mn> 
      </msup> 
     </mrow> 
    </math> estimate was affected by outliers, resulting in overestimation, particularly in small samples. The author <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref> recommended removing them before conducting MLR analysis. In case 1, reference <xref ref-type="bibr" rid="scirp.143538-2">
     [2]
    </xref> used a standardized residual of ±2. Outliers were retained within this interval, impacting the model’s validation parameters. The reviewed articles employed MLR training and testing on a single data distribution without considering sampling dispersion. According to <xref ref-type="bibr" rid="scirp.143538-8">
     [8]
    </xref>, Leave-One-Out Cross-Validation (LOO-CV) may introduce a bias selection and reduce model variance in small, unrepresentative samples, using nearly all data in each iteration. These samples can negatively affect its accuracy and performance. Methodological issues can produce unreliable models. Validating their adherence to MLR premises can statistically support them <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.143538-9">
     [9]
    </xref> <xref ref-type="bibr" rid="scirp.143538-10">
     [10]
    </xref>. This article analyzed five models from three studies utilizing their secondary data to address the following question:</p>
   <p>How valid were models generated using the QSAR approach?</p>
   <p>The research proposed that these models may not meet the premises of MLR due to the small sample sizes, which could affect their validity. This paper aimed to examine the violations or compliance with these prerequisites. It also targeted the consideration of the pros and cons of adopting them. The article analyzed assumptions and methods, assessing model conformity. Internal validity was evaluated by R<sup>2</sup> performance and RMSTr accuracy. External one shows predictive power in forecasting new molecular activities. A model is statistically valid if it significantly enhances these two indicators. This paper comprises five sections: premises of an MLR, materials and methods, results and discussion, and conclusion.</p>
  </sec><sec id="s2">
   <title>2. Premises of Multiple Linear Regression</title>
   <p>The models are required to meet the MLR guidelines, which encompass the relationships between DV and IVs. The quality of the database is analyzed as referenced by <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.143538-7">
     [7]
    </xref>. Cook’s Distance was employed to identify outliers. Its value has to be above 0.5 <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref>. Given the small sample sizes (N) of the five models, a threshold of 0.5 is considered more suitable than 4/N, as the latter results in higher thresholds. Both thresholds are conventional. It’s suggested to assess data homogeneity by using the Coefficient of Variation (CV). This latter is the ratio of the standard deviation to the mean. As noted by <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> and <xref ref-type="bibr" rid="scirp.143538-7">
     [7]
    </xref>, data are classified as uniform if the CV is less than 15%. According to <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref>, it’s necessary to establish the continuity of DV and IV, ensuring that they’re quantitative and not subject to any constraints. The independence of the DV is demonstrated by considering that data came from different molecules. The normality of its distribution is verified by employing the Kolmogorov-Smirnov (KS) test, as documented by <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.143538-9">
     [9]
    </xref>. IV variances are checked to guarantee they aren’t zero. The overall linearity must be examined.</p>
   <p>Reference <xref ref-type="bibr" rid="scirp.143538-10">
     [10]
    </xref> suggests using the Loess line and scatter plots of residuals RES1 and RES2 to verify linear relationships between a DV and its IVs. RES1 represents model error term, while RES2 is derived from another, treating the predictor as a DV. The Loess curve is included in the graph. Linearity is confirmed when it aligns closely with the regression line. The absence of multicollinearity between the IVs must be demonstrated. The Variance Inflation Factor (VIF) makes it possible <xref ref-type="bibr" rid="scirp.143538-9">
     [9]
    </xref>. According to <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-9">
     [9]
    </xref> <xref ref-type="bibr" rid="scirp.143538-11">
     [11]
    </xref>, it must be less than 10. The sample size, N, is compared to its theoretical value using inequalities 1 and 2 as described by <xref ref-type="bibr" rid="scirp.143538-3">
     [3]
    </xref> and <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref> respectively.</p>
   <p>
    <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        N 
      </mi> 
      <mo> 
      </mo> 
      <mo>
        ≥ 
      </mo> 
      <mo> 
      </mo> 
      <mn>
        50 
      </mn> 
      <mo>
        + 
      </mo> 
      <mn>
        8 
      </mn> 
      <mtext>
          
      </mtext> 
      <mi>
        v 
      </mi> 
      <mi>
        i 
      </mi> 
     </mrow> 
    </math> (1)</p>
   <p>
    <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
      <mi>
        N 
      </mi> 
      <mo>
        ≥ 
      </mo> 
      <mn>
        104 
      </mn> 
      <mo>
        + 
      </mo> 
      <mi>
        v 
      </mi> 
      <mi>
        i 
      </mi> 
     </mrow> 
    </math> (2)</p>
   <p>vi represents the number of IVs.</p>
   <p>Three key assumptions about residuals need to be demonstrated.</p>
  </sec><sec id="s3">
   <title>3. Materials and Method</title>
   <p>The research examined three articles that forecasted molecule activity in combating serious diseases, treating each as an individual case.</p>
   <sec id="s3_1">
    <title>3.1. Data Sampling</title>
    <p>The procedures for analyzing the data were detailed comprehensively. Reference <xref ref-type="bibr" rid="scirp.143538-2">
      [2]
     </xref> aimed to predict outcomes using Model 1 with 12 molecules for MOLT3 cell lines and seven for HepG2 ones.</p>
    <p>Case 1</p>
    <p>The activity (IC50) values of N-Tosyl-1,2,3,4-Tetrahydroisoquinoline derivatives were measured. They utilized LOO-CV to create training and test sets, identifying outliers with standardized residual cut-offs of ±2. The modelling was performed with Weka <xref ref-type="bibr" rid="scirp.143538-12">
      [12]
     </xref>. Descriptors were computed using Gaussian 03 DFT/B3LYP/6-31(d) and Dragon, with redundant indicators removed by the Unsupervised Forward Selection algorithm. Key predictors were pinpointed employing stepwise SPSS analysis, and the main findings were summarized. Two models were involved in Case 1. The expression for Model 1 was</p>
    <p>MOLT3 = 2.01312*Mor32u + 64.533*Gu − 12.2097 (3)</p>
    <p>Model 2’s wording was</p>
    <p>HepG 2 = −892.215*PJI3 + 1.1322*Mor32u − 1.048 3*Mor31v + 6.6454 (4)</p>
    <p>Gu referred to the symmetric index, Mor32u, Mor31v, and Mor32u were 3D-MoRSE indexes, and PJI3 was the Petitjean 3D index. Case 2 aimed to predict the antiproliferative activity of 13 5-Arylidene Rhodamines and compare descriptors.</p>
    <p>Case 2</p>
    <p>Nine molecules were modelled for inhibiting human lung tumours (NCI-H727) and ductal carcinoma (MDA-MB-231), with four used for testing. MLR was performed with Excel and Gaussian 09. It generated IVs at the DFT/B3LYP/6-31 G(d) level. Model 3 for the MDA-MB-231 cell line had the expression of:</p>
    <p>ActMDA (PCI50) = −459.091 76 + 1.892 18*ELUMO + 0.08176*ν<sub>C</sub><sub>=O</sub> + 39.433 17*d_C-N (5)</p>
    <p>Model 4 for the NCI-H727 cell line was documented as follows:</p>
    <p>ActNCI (PCI50) = −612.154 55 + 2.145 18*ELUMO + 0.11987*ν<sub>C=O</sub> + 30.62007*d_C-N (6)</p>
    <p>ELUMO represented the lowest unoccupied molecular orbital energy, ν<sub>C</sub><sub>=O</sub> indicated the CO frequency of the five-membered ring, and d_C-N referred to its CN distance. The third study analyzed 14 compounds related to Thiazoline and Thiazine for their antitumour properties.</p>
    <p>Case 3</p>
    <p>The case 3 predicted their pulmonary effects on A-549 cells. Model 5 used 10 molecules for training and four for testing with Gaussian 09 at DFT/B3LYP/6-31+G(d, p) levels. MLR in Excel calculated IVs from their expressions. Model 5 from this case was described as follows:</p>
    <p>PIC50 = 2.26432 − 0.77981*mu + 0.44572*LogP (7)</p>
    <p>LogP measured lipophilicity, and mu represented the molecular dipole moment.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Data Analysis</title>
    <p>The research undertaken by <xref ref-type="bibr" rid="scirp.143538-2">
      [2]
     </xref> <xref ref-type="bibr" rid="scirp.143538-5">
      [5]
     </xref> <xref ref-type="bibr" rid="scirp.143538-6">
      [6]
     </xref> supplied secondary data concerning both independent and dependent variables. These are organized in the Appendix across <xref ref-type="table" rid="tableA1">
      Table A1
     </xref>, <xref ref-type="table" rid="tableA2">
      Table A2
     </xref>, <xref ref-type="table" rid="tableA3">
      Table A3
     </xref> and <xref ref-type="table" rid="tableA4">
      Table A4
     </xref>. They were employed to verify that Models 1 to 5 adhered to MLR premises. Their compliance with this latter was conducted using SPSS Statistics version 27.</p>
    <p>Cook’s distance identified outliers utilizing Analysis &gt; Regression &gt; Linear &gt; Save, then selecting “Cooks” in the “Distance” box. Calculate CV with Descriptive Analyze &gt; Descriptives to find the mean and standard deviation. The KS test assessed normality of DV via Analyze &gt; Nonparametric &gt; Univariate tests &gt; KS test, with the null hypothesis assuming its normal distribution. Multicollinearity was checked by employing VIF: Analyze &gt; Regression &gt; Linear &gt; Statistics &gt; Collinearity Diagnostics. For linearity, a scatter plot of predicted values (x-axis) and residuals (y-axis) was generated: Analysis &gt; Regression &gt; Linear &gt; Graphs, plotting ZRESID vs. ZPRED. In graph editor, select Elements, Fit Curve to Total, Loess; a curve oscillating around Y = 0 indicates linearity. The sample size was verified by comparing the molecules’ quantity exploited to the estimates in Equations (1) <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref> and (2) <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. Homoscedasticity of error term was examined as per <xref ref-type="bibr" rid="scirp.143538-12">
      [12]
     </xref>. The process included standard regression analysis, saving non-standardized predicted values and residuals, and performing the test. Unstandardized residuals were squared and utilized in a subsequent regression model. Achieve another one with RES_squared as the DV employing the same IVs. Realize the Breusch-Pagan test via: Analyze &gt; Regression &gt; Linear &gt; Save, ensuring that Unstandardized for Predicted Values and Residuals boxes are checked, which will generate PRE1 and RES1 columns in Data View. The residuals were squared via the Transform menu: Compute Variable, naming RES_squared and entering RES1*RES1 as the formula. The p-value from the Breusch-Pagan test in the ANOVA table indicates heteroscedasticity if below 0.05. The normal distribution of residuals is assessed with a P-P graph using linear regression. The author <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> suggests configuring statistics by selecting: “Estimation,” “Hypothesis test,” and “Residual predictions.” After choosing “Histograms” and “Normal P-P Chart” in the chart menu, the PP plot is generated. Points near the diagonal suggest a normal distribution of residuals. The Durbin-Watson D. identifies error term independence <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. References <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> and <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> describe the standard regression procedure: select the “Statistics” option and check the “Durbin-Watson” box under “Residuals.” Durbin-Watson D. close to 2 indicates independent residuals, next to 0 shows positive autocorrelation, and around 4 proposed a negative one <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. MLR premise conformity analyses were performed for all models. The findings are presented and discussed.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. Results and Discussion</title>
   <p>Models 1 and 2 related to N-Tosyl-1,2,3,4-Tetrahydroisoquinoline were evaluated. According to <xref ref-type="bibr" rid="scirp.143538-1">
     [1]
    </xref> <xref ref-type="bibr" rid="scirp.143538-4">
     [4]
    </xref> <xref ref-type="bibr" rid="scirp.143538-7">
     [7]
    </xref>, DVs and IVs data must be free of outliers.</p>
   <sec id="s4_1">
    <title>4.1. Compliance with Data on N-Tosyl-Tetrahydroisoquinoline Compounds</title>
    <p>In Model 1, molecule 8 (4h) had a Cook’s distance of 0.507 <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref>. In Model 2, compound 4 (4h) had a statistic of 4.372, above 0.5, while molecule 6 (4k) had 0.812. These data were outliers <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref>. They were removed before linear regression analysis <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. Those for models 1 and 2 were heterogeneous, with CVs over 15% <xref ref-type="bibr" rid="scirp.143538-7">
      [7]
     </xref>. For Model 1, DV MOLT3 was 39%, IV Mor32u was 41%. In Model 2, DV HepG2 was 16%, IV Mor32u was 59%, and IV Mor31v was 19%.</p>
    <p>The DVs and IVs in models 1 and 2 were continuous; they were quantitative with no constraints <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. DV data were normally distributed (KS test: p = 0.730). MOLT3 and HepG2 activities originated from different molecules, ensuring their independence from each other. The non-zero variance premises were violated. Furthermore, Model 1 had VIFs of 1.028. Those of Model 2 were worth 1.899 for IV Mor32u, 1.872 for Mor31v and 1.030 for PIJ3. All values were below 10, indicating an absence of multicollinearity between the IVs of the two models <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-11">
      [11]
     </xref>. <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref> indicates that Model 1 has deviated from the y=0 line, thereby compromising its linearity <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-13">
      [13]
     </xref>. Similarly, <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>, associated with Model 2, depicts a comparable scenario.</p>
    <p>The models used small samples. Model 1 employed 11 molecules, less than the 66 and 106 obtained with Equations (1) <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref> and (2) <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. Likewise, Model 2 exploited compounds instead of the needed 74 and 107. The Breusch-Pagan test revealed heteroscedasticity in Model 1 (F = 0.117, p = 0.891) and Model 2 (F = 0.545, p = 0.665) <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-12">
      [12]
     </xref>.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. Model 1’s linear relationship analysis.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId28.jpeg?20250625023303" />
    </fig>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Model 2’s linear relationship analysis.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId29.jpeg?20250625023303" />
    </fig>
    <p>The points in the normal P-P plot of Model 1 didn’t align along the diagonal, as shown in <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref>. Similarly, those in Model 2, illustrated in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>, didn’t align either. In both cases, the residuals weren’t normally distributed around zero <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. Model 1 produced a Durbin-Watson statistic of 2.668, while Model 2 generated 2.736. These values suggest that errors in both models are independent <xref ref-type="bibr" rid="scirp.143538-8">
      [8]
     </xref>. A summary of the results can be found in <xref ref-type="table" rid="table1">
      Table 1
     </xref>. The models respected the same premises.</p>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. Assessment of Model 1 residuals for normal distribution.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId30.jpeg?20250625023304" />
    </fig>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>Figure 4. Assessment of Model 2 residuals for normal distribution.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId31.jpeg?20250625023303" />
    </fig>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143538-"></xref>Table 1. Analysis results: compliance summary for Models 1 and 2.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="acenter" width="19.38%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="44.84%" colspan="2"><p style="text-align:center">Model 1</p></td> 
       <td class="custom-bottom-td acenter" width="46.55%" colspan="2"><p style="text-align:center">Model 2</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="19.38%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="20.66%"><p style="text-align:center">Violated Premises</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="24.18%"><p style="text-align:center">Respected Premises</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="23.25%"><p style="text-align:center">Violated Premises</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="23.30%"><p style="text-align:center">Respected Premises</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="19.38%"><p style="text-align:center">Data homogeneity</p></td> 
       <td class="custom-top-td acenter" width="20.66%"><p style="text-align:center">Data homogeneity</p></td> 
       <td class="custom-top-td acenter" width="24.18%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="23.25%"><p style="text-align:center">Data homogeneity</p></td> 
       <td class="custom-top-td acenter" width="23.30%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.38%"><p style="text-align:center">DV</p></td> 
       <td class="acenter" width="20.66%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="24.18%"><p style="text-align:center">DV continuity</p><p style="text-align:center">DV normality</p><p style="text-align:center">DV independence</p></td> 
       <td class="acenter" width="23.25%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="23.30%"><p style="text-align:center">DV continuity</p><p style="text-align:center">DV normality</p><p style="text-align:center">DV independence</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.38%"><p style="text-align:center">IV </p></td> 
       <td class="acenter" width="20.66%"><p style="text-align:center">Non-zero variance</p></td> 
       <td class="acenter" width="24.18%"><p style="text-align:center">IV continuity</p><p style="text-align:center">Absence of multicollinearity</p></td> 
       <td class="acenter" width="23.25%"><p style="text-align:center">Non-zero variance</p></td> 
       <td class="acenter" width="23.30%"><p style="text-align:center">IV continuity</p><p style="text-align:center">Absence of multicollinearity </p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.38%"><p style="text-align:center">Relationship between DV and its IVs</p></td> 
       <td class="acenter" width="20.66%"><p style="text-align:center">Overall linearity</p></td> 
       <td class="acenter" width="24.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="23.25%"><p style="text-align:center">Overall linearity</p></td> 
       <td class="acenter" width="23.30%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.38%"><p style="text-align:center">Sample</p></td> 
       <td class="acenter" width="20.66%"><p style="text-align:center">Sample size</p></td> 
       <td class="acenter" width="24.18%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="23.25%"><p style="text-align:center">Sample size</p></td> 
       <td class="acenter" width="23.30%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="19.38%"><p style="text-align:center">Residues’ characteristics</p></td> 
       <td class="acenter" width="20.66%"><p style="text-align:center">Residues’ homoscedasticity </p><p style="text-align:center">Normality of residual distribution</p></td> 
       <td class="acenter" width="24.18%"><p style="text-align:center">Residues’ Independence</p></td> 
       <td class="acenter" width="23.25%"><p style="text-align:center">Residues’ homoscedasticity Normality of residual distribution </p></td> 
       <td class="acenter" width="23.30%"><p style="text-align:center">Residues’ independence</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The two models followed the assumptions of continuity, normal distribution, and independence concerning the DVs. They adhered to the continuity of the IVs and avoided multicollinearity. They violated basic hypotheses related to non-zero variances of IVs, overall linearity of the model, and adequate sample size. They also transgressed premises regarding homoscedasticity and independence of the error term. The effectiveness of Models 1 and 2 is determined by their capacity to accurately predict the cytotoxic activity of N-Tosyl-Tetrahydroisoquinoline molecules.</p>
    <p>Respect for the premises underlines this possibility. Without multicollinearity, the relationship between the DV and its IVs can be pinpointed. This precisely identifies the β<sub>X</sub> coefficients and accounts for variations in the IVs. In concert with <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>, the DVs MOLT3 and HepG2 can be exactly predicted due to their continuity, normal distribution, and independence. Independent residuals ensure accurate confidence intervals <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-14">
      [14]
     </xref> and parameter estimates for both models <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-14">
      [14]
     </xref>, this improves the reliability of their outcomes and forecasts. Evaluating confidence intervals correctly avoids biases in estimations and predictions. Deviation from premises reduces the projection precision of the two models.</p>
    <p>The nonlinearity impacts β<sub>X</sub> estimates, making it difficult to understand the relationship between HepG2 and its predictors PIJ3, Mor32u, and Mor31v, leading to unreliable projections <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-15">
      [15]
     </xref>. A small sample size may result in incorrect R<sup>2</sup> and β<sub>X</sub> coefficients <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> <xref ref-type="bibr" rid="scirp.143538-15">
      [15]
     </xref> <xref ref-type="bibr" rid="scirp.143538-16">
      [16]
     </xref>. It raises the variance of estimators, complicating the detection of IV effects <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>, heteroscedasticity affects hypothesis testing and confidence intervals by misestimating variances, resulting in incorrect β<sub>X</sub> coefficients. Non-normally distributed residuals increase the risk of type I errors, leading to inaccuracies in model predictions <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>.</p>
    <p>Violating global linearity, homoscedasticity, and normal distribution of residuals reduces the precision of projections from Models 1 and 2. Respecting multicollinearity doesn’t sufficiently compensate for these issues. The quality of each model is also evaluated based on their ability to generalize results.</p>
    <p>The small sample size increases sensitivity to fluctuations and lower external validity, making it harder to generalize predictions to additional N-Tosyl-Tetrahydroisoquinoline molecules <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. The residuals ’heteroscedasticity can lead to unstable MOLT3 and HepG2 predictions, causing incorrect conclusions about IVs variations <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. However, their independence ensures that each model is robust and applicable to new derivatives from N-Tosyl-Tetrahydroisoquinoline compounds <xref ref-type="bibr" rid="scirp.143538-17">
      [17]
     </xref>. Generalizing the outcomes of both models is difficult despite the contribution of their quality. Models 1 and 2 don’t meet the premises, allowing the activity of the molecules studied to be accurately predicted and generalized to others in the same family. Consequently, their validity remains uncertain. Additionally, Models 3 and 4 related to 5-Arylidene Rhodanine compounds were analyzed thoroughly.</p>
   </sec>
   <sec id="s4_2">
    <title>4.2. Compliance with Data on 5-Arylidene Rhodanine Compounds</title>
    <p>In Model 3, the DV ActMDA score for molecule 4 was an outlier with a Cook’s distance of 0.662, exceeding 0.5. It was removed <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. Cook distances in Model 4 ranged from 0.007 to 0.236, all below 0.5. No outlier was found in DV ActNCI <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref>. Furthermore, IVs data were homogeneous for each model <xref ref-type="bibr" rid="scirp.143538-7">
      [7]
     </xref>. The CV of IV ELUMO was 8%, and those of IV d_CN and νC=O were identical at 0.2%. The DVs data were heterogeneous <xref ref-type="bibr" rid="scirp.143538-7">
      [7]
     </xref>, with CVs of 57.6% for DV ActNCI and 58.8% for DV ActMDA.</p>
    <fig id="fig5" position="float">
     <label>Figure 5</label>
     <caption>
      <title>Figure 5. Model 3’s linear relationship analysis.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId32.jpeg?20250625023306" />
    </fig>
    <fig id="fig6" position="float">
     <label>Figure 6</label>
     <caption>
      <title>Figure 6. Model 4’s linear relationship analysis.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId33.jpeg?20250625023306" />
    </fig>
    <p>The DVs data of both models exhibited a normal distribution (p = 0.200) <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. They were continuous as they were quantitative and didn’t face any limitations, such as the IVs ELUMO, ν<sub>C</sub><sub>=O</sub> and d_CN <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. The DVs data originated from different 5-Arylidene Rhodanine compounds that were independent. In Models 3 and 4, ELUMO and ν<sub>C</sub><sub>=O</sub> displayed non-zero variances, while d_CN presented zero variance. Model 3 indicated VIFs of 1.134, and Model 4 had VIFs of 1.077, 1.659, and 1.63, all below 10, suggesting no multicollinearity <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-11">
      [11]
     </xref>. <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref> shows Model 3’s Loess line deviating significantly from zero. <xref ref-type="fig" rid="fig6">
      Figure 6
     </xref> depicts Model 4’s deviation. The condition for overall linearity wasn’t met <xref ref-type="bibr" rid="scirp.143538-13">
      [13]
     </xref>. The sample sizes for the two models were insufficient, with 13 compounds instead of the required 74 or 107 as outlined by Equations (1) <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref> and (2) <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>.</p>
    <fig id="fig7" position="float">
     <label>Figure 7</label>
     <caption>
      <title>Figure 7. Assessment of Model 3 residuals for normal distribution.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId34.jpeg?20250625023306" />
    </fig>
    <fig id="fig8" position="float">
     <label>Figure 8</label>
     <caption>
      <title>Figure 8. Assessment of Model 4 residuals for normal distribution.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId35.jpeg?20250625023306" />
    </fig>
    <p>The Breusch-Pagan distribution tests showed no significant results, with a Fisher coefficient of F = 1.049 (p = 0.418) for Model 3 and F = 0.284 (p = 0.836) for Model 4. Neither model met the assumption of residuals’ homoscedasticity <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> <xref ref-type="bibr" rid="scirp.143538-12">
      [12]
     </xref>. <xref ref-type="fig" rid="fig7">
      Figure 7
     </xref>, related to Model 3, demonstrates the non-alignment of the points on the P-P graph with the diagonal. Similarly, <xref ref-type="fig" rid="fig8">
      Figure 8
     </xref>, associated with Model 4, shows the same pattern. Consequently, under these conditions, the errors for DV ActMDA and ActNCI weren’t normally distributed <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>, the Durban-Watson D. suggested positive autocorrelation (ActMDA: 1.184; ActNCI: 0.997). The residues for models 3 and 4 weren’t independent. A summary of the results can be found in <xref ref-type="table" rid="table2">
      Table 2
     </xref>.</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143538-"></xref>Table 2. Analysis results: compliance summary for Models 3 and 4.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="acenter" width="17.72%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="40.36%" colspan="2"><p style="text-align:center">Model 3</p></td> 
       <td class="custom-bottom-td acenter" width="41.92%" colspan="2"><p style="text-align:center">Model 4</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="17.72%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="20.90%"><p style="text-align:center">Violated Premises</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.46%"><p style="text-align:center">Respected Premises</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="21.20%"><p style="text-align:center">Violated Premises</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="20.72%"><p style="text-align:center">Respected Premises</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="17.72%"><p style="text-align:center">Data homogeneity</p></td> 
       <td class="custom-top-td acenter" width="20.90%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="19.46%"><p style="text-align:center">Data Homogeneity </p></td> 
       <td class="custom-top-td acenter" width="21.20%"><p style="text-align:center"></p></td> 
       <td class="custom-top-td acenter" width="20.72%"><p style="text-align:center">Data Homogeneity</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.72%"><p style="text-align:center">DV</p></td> 
       <td class="acenter" width="20.90%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="19.46%"><p style="text-align:center">DV continuity</p><p style="text-align:center">DV normality</p><p style="text-align:center">DV independence</p></td> 
       <td class="acenter" width="21.20%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="20.72%"><p style="text-align:center">DV ActMDA continuity</p><p style="text-align:center">DV ActMDA normality</p><p style="text-align:center">DV ActMDA independence</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.72%"><p style="text-align:center">IV </p></td> 
       <td class="acenter" width="20.90%"><p style="text-align:center">Non-zero variance</p></td> 
       <td class="acenter" width="19.46%"><p style="text-align:center">IV continuity</p><p style="text-align:center">Absence of Multicollinearity</p></td> 
       <td class="acenter" width="21.20%"><p style="text-align:center">Non-zero variance</p></td> 
       <td class="acenter" width="20.72%"><p style="text-align:center">IV continuity </p><p style="text-align:center">Absence of Multicollinearity</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.72%"><p style="text-align:center">Linear Relationship </p></td> 
       <td class="acenter" width="20.90%"><p style="text-align:center">Overall linearity</p></td> 
       <td class="acenter" width="19.46%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="21.20%"><p style="text-align:center">Overall linearity </p></td> 
       <td class="acenter" width="20.72%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.72%"><p style="text-align:center">Sample</p></td> 
       <td class="acenter" width="20.90%"><p style="text-align:center">Sample size</p></td> 
       <td class="acenter" width="19.46%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="21.20%"><p style="text-align:center">Sample size </p></td> 
       <td class="acenter" width="20.72%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="17.72%"><p style="text-align:center">Residues’ characteristics</p></td> 
       <td class="acenter" width="20.90%"><p style="text-align:center">Residues’ homoscedasticity </p><p style="text-align:center">Normality of Residuals Distribution</p><p style="text-align:center">Residues’ independence</p></td> 
       <td class="acenter" width="19.46%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="21.20%"><p style="text-align:center">Residues’ homoscedasticity</p><p style="text-align:center">Normality of Residuals Distribution</p><p style="text-align:center">Residues independence</p></td> 
       <td class="acenter" width="20.72%"><p style="text-align:center"></p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>Models 3 and 4 met data homogeneity criteria, ensuring continuity and avoiding multicollinearity of the IVs. However, they fall to respect non-zero variances, linearity, adequate sample size, homoscedasticity, normal distribution, or independence of residual requirements. Their effectiveness is linked to their ability to accurately predict the antiproliferative activity of 5-Arylidene Rhodanine molecules. Compliance with certain assumptions underlying MLR predisposes them.</p>
    <p>As noted by references <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> and <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>, the normal distribution and independence of DV data in both models enable precise estimation of molecular activities, enhancing consistency between dependent and independent variables. Continuity in IVs ELUMO, ν<sub>C</sub><sub>=O</sub>, and d_CN achieves similar outcomes. The absence of multicollinearity simplifies models 3 and 4 <xref ref-type="bibr" rid="scirp.143538-18">
      [18]
     </xref>. However, predicting the antiproliferative effects of 5-Arylidene Rhodanines remains challenging due to certain premise violations.</p>
    <p>The non-zero variance of IVs (ELUMO, freq. CO, and d_CN) limits models 3 and 4 in assessing the accuracy of relationships between variables. Moreover, the absence of linearity can cause unclear predictions for DVs ActMDA and ActNCI <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. As highlighted in references <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-15">
      [15]
     </xref>, it introduces bias into β<sub>X</sub> coefficients. This lack of precision doesn’t significantly improve the accuracy of forecasts. Consequently, parameter estimates become incorrect, resulting in unreliable values for IV ELUMO, ν<sub>C</sub><sub>=O</sub>, and d_CN, which increases the risk of type II errors <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref>. Thus, the determination of R<sup>2</sup> and β<sub>X</sub> remains imprecise <xref ref-type="bibr" rid="scirp.143538-15">
      [15]
     </xref>. Heteroscedasticity reduces the variance of ActMDA and ActNCI, impacting hypothesis tests and confidence intervals similarly to β<sub>X</sub> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. It may lead to false conclusions regarding the effects of IVs ELUMO, ν<sub>C</sub><sub>=O</sub>, and d_CN <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-14">
      [14]
     </xref>, error dependence results in incorrect parameter estimates and affects significance tests and confidence intervals, indicating unincorporated structures in both models. Furthermore, non-normally distributed residuals elevate the likelihood of type I errors, resulting in model prediction inaccuracies <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>.</p>
    <p>Models 3 and 4’s prediction accuracy is reduced due to violations of linearity, homoscedasticity, normal distribution, and independence of residuals. Additionally, their effectiveness also depends on the generalizability of the results.</p>
    <p>The small sample size prevents extrapolating the results of the two models to another 5-Arylidene Rhodanine molecule <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref>. For <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>, Heteroscedasticity may introduce instability in the effects of the IVs ELUMO, ν<sub>C</sub><sub>=O</sub>, and d_CN. Models 3 and 4 can’t measure variations in VI. Moreover, the error term dependence can decrease performance and restrict the generalizability of results to new molecules from 5-Arylidene Rhodanine <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-17">
      [17]
     </xref>. Additionally, these two models don’t accurately predict their antiproliferative activity. Hence, it’s advised that they be re-evaluated. The study further examined Model 5 in relation to Thiazine and Thiazoline compounds.</p>
   </sec>
   <sec id="s4_3">
    <title>4.3. Compliance with Data Premises on Thiazoline and Thiazine Compounds</title>
    <p>The cook’s distance for outliers ranged from 0.003 to 0.346, all below 0.5. Model 5 had none <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref>. IV mu data were uniform (CV 3%), while those of DV PIC50 and IV logP weren’t (CVs 23% and 30%) <xref ref-type="bibr" rid="scirp.143538-7">
      [7]
     </xref>. DV PIC50 and its IVs were continuous <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>; their data were quantitative and unrestricted. The KS test confirmed the DVPIC50’s normal distribution (p = 0.328) <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. Data from distinct Thiazoline and Thiazine compounds were independent. Variances were 0.013 for IV mu and 0.450 for IV LogP, meeting non-nullity assumption. VIFs were 1.072, indicating no multicollinearity <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-11">
      [11]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-13">
      [13]
     </xref>, the Loess curve approached zero. <xref ref-type="fig" rid="fig9">
      Figure 9
     </xref> illustrates that the model exhibited overall linearity <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> <xref ref-type="bibr" rid="scirp.143538-13">
      [13]
     </xref>. With only 14 compounds studied, the sample size was insufficient as Equations (1) <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref> and (2) <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> required 66 or 106 molecules. The Breush-Pagan test wasn’t significant (F = 1.546; p = 0.256), confirming the heteroscedasticity of this model <xref ref-type="bibr" rid="scirp.143538-12">
      [12]
     </xref>. The normal P-P plot in <xref ref-type="fig" rid="fig10">
      Figure 10
     </xref> shows the Model 5 error distribution, but the points didn’t align with the diagonal.</p>
    <fig id="fig9" position="float">
     <label>Figure 9</label>
     <caption>
      <title>Figure 9. Model 5’s linear relationship analysis.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId36.jpeg?20250625023308" />
    </fig>
    <fig id="fig10" position="float">
     <label>Figure 10</label>
     <caption>
      <title>Figure 10. Assessment of Model 5 residuals for normal distribution.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1710190-rId37.jpeg?20250625023308" />
    </fig>
    <p>Model 5 residues weren’t normally distributed around zero <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>, although they were independent <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>; Durbin D. value was 1.916, indicating no autocorrelation <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. The primary findings of this analysis are presented in <xref ref-type="table" rid="table3">
      Table 3
     </xref>.</p>
    <p>Model 5 met the DV PCI50 criteria for continuity, independence, and normality. It adheres to IVs continuity without multicollinearity issues. Although linearity and normally distributed error terms were followed, data homogeneity, sample size, and residuals’ homoscedasticity were inadequate. Model 5’s prediction accuracy for molecules’ antitumour activity depends on specific premises.</p>
    <p>The continuity of the DV PIC50 and its IVs mu and LogP, along with non-zero variances and the absence of multicollinearity, facilitate a clearer understanding of the structural relationships between them <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> <xref ref-type="bibr" rid="scirp.143538-18">
      [18]
     </xref>. Non-zero variances and no multicollinearity ensure small confidence intervals and reliable significance tests <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-18">
      [18]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-10">
      [10]
     </xref>, Model 5’s linearity allows for appropriate correlations between variables and valid confidence intervals, which enhances forecast accuracy. Independent residuals result in precise parameter estimates <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref> and confidence intervals <xref ref-type="bibr" rid="scirp.143538-14">
      [14]
     </xref>. This improves the reliability of its outcomes and forecasts. Correctly evaluated confidence intervals avoid biases in parameter estimation and predictions <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref>. Nonetheless, violating several premises can restrict the model’s precision.</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143538-"></xref>Table 3. Analysis results: compliance summary for Model 5.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="30.22%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="27.58%"><p style="text-align:center">Violated Premises</p></td> 
       <td class="custom-bottom-td acenter" width="42.20%"><p style="text-align:center">Respected Premises</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="30.22%"><p style="text-align:center">Data homogeneity</p></td> 
       <td class="custom-top-td acenter" width="27.58%"><p style="text-align:center">Data homogeneity</p></td> 
       <td class="custom-top-td acenter" width="42.20%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">DV characteristics</p></td> 
       <td class="acenter" width="27.58%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="42.20%"><p style="text-align:center">DV PIC50 continuity</p><p style="text-align:center">DV PIC50 normality</p><p style="text-align:center">DV PIC50 independence</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">IV characteristics</p></td> 
       <td class="acenter" width="27.58%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="42.20%"><p style="text-align:center">IV mu and LogP continuity</p><p style="text-align:center">Absence of multicollinearity</p><p style="text-align:center">Non-zero variance</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">Linear Relationship</p></td> 
       <td class="acenter" width="27.58%"><p style="text-align:center"></p></td> 
       <td class="acenter" width="42.20%"><p style="text-align:center">Overall linearity</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">Sample </p></td> 
       <td class="acenter" width="27.58%"><p style="text-align:center">Sample size</p></td> 
       <td class="acenter" width="42.20%"><p style="text-align:center"></p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">Residues’ characteristics</p></td> 
       <td class="acenter" width="27.58%"><p style="text-align:center">Residues’ homoscedasticity</p></td> 
       <td class="acenter" width="42.20%"><p style="text-align:center">Normality of residual distribution Residues’ independence</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The limited sample size compromises the reliability of parameter estimators, rendering them inadequate for accurately representing Thiazoline or Thiazine activity <xref ref-type="bibr" rid="scirp.143538-19">
      [19]
     </xref>. According to <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>, error term’s heteroscedasticity affects the correct understanding of the effects of IVs LogP and mu. It can misestimate the variance of DV PIC50. This results in less precise βx coefficients.</p>
    <p>The non-normality of the residual distribution impacts reliability and invalidates confidence intervals and hypothesis tests <xref ref-type="bibr" rid="scirp.143538-1">
      [1]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref> <xref ref-type="bibr" rid="scirp.143538-9">
      [9]
     </xref>. Model 5’s validation evaluates outcome generalizability, but the insignificant sample size may influence reliability. Changes in DV PIC50 and IVs mu and LogP increase sensitivity, affecting adaptability and predictive accuracy <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref>. Results vary by Thiazoline and Thiazine studied. On the other hand, the independence of residuals makes Model 5 robust and applicable to different datasets <xref ref-type="bibr" rid="scirp.143538-17">
      [17]
     </xref>. The small sample size remains a significant concern, which hinders the generalization of the findings of Model 5 to additional molecules from Thiazine or Thiazolines <xref ref-type="bibr" rid="scirp.143538-3">
      [3]
     </xref> <xref ref-type="bibr" rid="scirp.143538-4">
      [4]
     </xref>. This limitation invalidates it. Model 5 excels in forecasting the pulmonary effects of Thiazine or Thiazoline on A-549 cells. However, it struggles to apply this prediction to similar molecules. Increasing the sample size could corroborate it. <xref ref-type="table" rid="table4">
      Table 4
     </xref> outlines how the five models conform to MLR premises. It offers a perspective to conclude the article.</p>
    <table-wrap id="table4">
     <label>
      <xref ref-type="table" rid="table4">
       Table 4
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.143538-"></xref>Table 4. Summary of compliance analysis for all models.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="18.92%"><p style="text-align:center"></p></td> 
       <td class="custom-bottom-td acenter" width="30.22%"><p style="text-align:center">MLR Premises</p></td> 
       <td class="custom-bottom-td acenter" width="10.16%"><p style="text-align:center">Model 1</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Model 2</p></td> 
       <td class="custom-bottom-td acenter" width="10.16%"><p style="text-align:center">Model 3</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Model 4</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Model 5</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="18.92%"><p style="text-align:center">Data set </p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="30.22%"><p style="text-align:center">The data are homogeneous</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
      </tr> 
      <tr> 
       <td rowspan="3" class="custom-top-td acenter" width="18.92%"><p style="text-align:center">DV</p></td> 
       <td class="custom-top-td acenter" width="30.22%"><p style="text-align:center">DV is continuous</p></td> 
       <td class="custom-top-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">The DV follows a normal distribution.</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="30.22%"><p style="text-align:center">DV data are independent</p></td> 
       <td class="custom-bottom-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td rowspan="3" class="custom-top-td acenter" width="18.92%"><p style="text-align:center">IV </p></td> 
       <td class="custom-top-td acenter" width="30.22%"><p style="text-align:center">IVs are continuous</p></td> 
       <td class="custom-top-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">The IVs exhibit non-zero variances</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="30.22%"><p style="text-align:center">No multicollinearity between IVs</p></td> 
       <td class="custom-bottom-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="custom-bottom-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="18.92%"><p style="text-align:center">Relationship between the DV and its IVs</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="30.22%"><p style="text-align:center">The DV-IV relationship is usually linear.</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="18.92%"><p style="text-align:center">Sample</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="30.22%"><p style="text-align:center">The sample size satisfies Green’s or Pallant’s conditions.</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
      </tr> 
      <tr> 
       <td rowspan="3" class="custom-top-td acenter" width="18.92%"><p style="text-align:center">Residues’ characteristics</p></td> 
       <td class="custom-top-td acenter" width="30.22%"><p style="text-align:center">Residues’ homoscedasticity </p></td> 
       <td class="custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-top-td acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="custom-top-td acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">Residuals are normally distributed.</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="30.22%"><p style="text-align:center">The residues are independent</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Respected</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
       <td class="acenter" width="10.16%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Violated</p></td> 
       <td class="acenter" width="10.18%"><p style="text-align:center">Respected</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
  </sec><sec id="s5">
   <title>5. Conclusions</title>
   <p>This research assessed the compliance of five QSAR models with MLR premises, uncovering limitations due to validation criteria and sample size. Assumptions were summarized, and conformity was analyzed by comparing data characteristics to MLR premises. The accuracy of its predictions and the possibility of generalizing its results were emphasized. Models 1, 2, and 3 excluded outliers to meet MLR requirements. Homogeneity was proven in models 3 and 4. While all adhered to the basic hypotheses for DVs, they partially respected those of IVs. Except Model 5, others violated non-zero variance assumptions for IVs. All models complied with the premise that IVs shouldn’t exhibit multicollinearity. They failed to satisfy sufficient sample size criteria. They breached the principles relating to homoscedasticity of residuals, but models 1, 2 and 5 conform to those associated with their independence. Model 5 demonstrated global linearity with an error term that was normally distributed.</p>
   <p>Adherence to MLR assumptions enhances forecast accuracy and generalizability. Conversely, violations of these assumptions reduce their effectiveness. Factors such as sample size, residuals’ heteroscedasticity, and interdependence limit generalizability. Models 1 to 4 face issues with nonlinearity and non-normal residuals, while Models 3 and 4 don’t address the error term dependence. Despite challenges in validation due to a small sample size, Model 5 accurately predicted the pulmonary effects of Thiazine or Thiazoline on A-549 cells. Verify that the data matches MLR premises when using these projected models.</p>
  </sec><sec id="s6">
   <title>Appendix. Pooled Training and Test Data</title>
   <table-wrap id="table5">
    <label>
     <xref ref-type="table" rid="table5">
      Table 5
     </xref></label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.143538-"></xref><p class="imgGroupCss_v"><img class=" imgMarkCss lazy" data-original="https://html.scirp.org/file/1710190-rId49.jpeg?20250625023312" /></p></title>
    </caption>
   </table-wrap>
   <p>
    <xref ref-type="bibr" rid="scirp.143538-"></xref><p class="imgGroupCss_v"><img class=" imgMarkCss lazy" data-original="https://html.scirp.org/file/1710190-rId54.jpeg?20250625023313" /></p></p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.143538-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Field, A. (2009) Discovering Statistics Using SPSS. Sage Publication Ltd.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pingaew, R., Worachartcheewan, A., Nantasenamat, C., Prachayasittikul, S., Ruchirawat, S. and Prachayasittikul, V. (2013) Synthesis, Cytotoxicity and QSAR Study of N-tosyl-1,2,3,4-tetrahydroisoquinoline Derivatives. Archives of Pharmacal Research, 36, 1066-1077. &gt;https://doi.org/10.1007/s12272-013-0111-9
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pallant, J. (2023) SPSS Survival Manual: A Step-by-Step Guide to Data Analysis Using IBM SPSS. 7th Edition, Open University Press.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Green, S.B. (1991) How Many Subjects Does It Take to Do a Regression Analysis. Multivariate Behavioral Research, 26, 499-510. &gt;https://doi.org/10.1207/s15327906mbr2603_7
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Coulibaly, W.K., Affi, S.T., James, T., Koné, M.G.-R., Yao, A.E.B., Dago, C.D., et al. (2022) Anti-Proliferative Activity Study on 5-Arylidene Rhodanine Derivatives Using Density Functional Theory (DFT) and Quantitative Structure Activity Relationship (QSAR). International Journal of Computational and Theoretical Chemistry, 10, 1-8.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Dembelé, G.S., Tuo, N.T., Konaté, F., Soro, D., Konaté, B. and Ziao, N. (2022) Quantitative Structure Activity Relationship (QSAR) Study of a Series of Molecules Derived from Thiazoline and Thiazine Multithioether Having Activity against Antitumor Activity (A-549). International Journal of Chemical and Life Sciences, 11, 2426-2435. &gt;https://www.researchgate.net/publication/364965605_Quantative_Structure_Activity_Relationship_QSAR_Study_of_a_Series_of_Molecules_Derived_from_Thiazoline_and_Thiazine_Multithioether_Having_Activity_against_Antitumor_Activity_A-549
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Baillargeon, G. (2010) Méthodologies et techniques statistiques, Trois-Rivières: Bibliothèque nationale du Québec, SMG.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lv, L., Song, X. and Sun, W. (2020) Modify Leave-One-Out Cross Validation by Moving Validation Samples around Random Normal Distributions: Move-One-Away Cross Validation. Applied Sciences, 10, Article No. 2448. &gt;https://doi.org/10.3390/app10072448
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Flatt, C. and Jacobs, R.L. (2019) Principle Assumptions of Regression Analysis: Testing, Techniques, and Statistical Reporting of Imperfect Data Sets. Advances in Developing Human Resources, 21, 484-502. &gt;https://doi.org/10.1177/1523422319869915
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Schmidt, A.F. and Finan, C. (2018) Linear Regression and the Normality Assumption. Journal of Clinical Epidemiology, 98, 146-151. &gt;https://doi.org/10.1016/j.jclinepi.2017.12.006
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yassine, T. (2020) Contribution des technologies de l’information et de la communication au succès de la collaboration client-fournisseur en développement de produits nouveaux. Université de Grenoble Alpes. &gt;https://theses.hal.science/tel-02931916/
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Johnson, R. and Wichern, D. (2018) Applied Multivariate Statistical Analysis. 6th Edition, Pearson.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nachhilfe, S. (2022) Comment vérifier la condition de linéarité pour le modèle de régression linéaire dans R et SPSS? &gt;https://statistiknachhilfe.ch/fr/2022/12/09/voraussetzung-lineare-regression-linearitat/ 
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Dodge, Y. and Rousson, V. (2004) Analyse de régression appliquée. Dunod.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Moore, D.S., McCabe, G.P. and Craig, B. (2021) Introduction to the Practice of Statistics. 10th Edition, W.H. Freeman.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Sing, V. (2025) Multicollinéarité dans la régression: Un guide pour les scientifiques des données. &gt;https://www.datacamp.com/fr/tutorial/multicollinearity?form=MG0AV3 
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Della-Vedova, C. (s.d.) Régression linéaire simple: Quand les hypothèses ne sont pas satisfaites. &gt;https://delladata.fr/regression-lineaire-simple-quand-les-hypotheses-ne-sont-pas-satisfaites/
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lind, A.D., Marchal, W., Mason, D.R., Satya, G.D., Santosh, K. and Singh, J. (2007) Méthodes statistiques pour les sciences de la gestion. Les éditions de la Chenelière Inc.
    </mixed-citation>
   </ref>
   <ref id="scirp.143538-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Draper, N.R. and Smith, H. (1998) Applied Regression Analysis. 3rd Edition, Wiley. &gt;https://doi.org/10.1002/9781118625590
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>