<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.4 20241031//EN" "JATS-journalpublishing1-4.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.4" xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">jdaip</journal-id>
      <journal-title-group>
        <journal-title>Journal of Data Analysis and Information Processing</journal-title>
      </journal-title-group>
      <issn pub-type="epub">2327-7203</issn>
      <issn pub-type="ppub">2327-7211</issn>
      <publisher>
        <publisher-name>Scientific Research Publishing</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4236/jdaip.2026.141004</article-id>
      <article-id pub-id-type="publisher-id">jdaip-148180</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
        <subj-group>
          <subject>Computer Science</subject>
          <subject>Communications</subject>
          <subject>Physics</subject>
          <subject>Mathematics</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Cumulative Link Modeling of Ordinal Outcomes in the National Health Interview Survey Data: Application to Depressive Symptom Severity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Williams</surname>
            <given-names>Andre</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name name-style="western">
            <surname>Louis</surname>
            <given-names>Louisana</given-names>
          </name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
      </contrib-group>
      <aff id="aff1"><label>1</label> Christine E. Lynn College of Nursing, Florida Atlantic University, Boca Raton, FL, USA </aff>
      <author-notes>
        <fn fn-type="conflict" id="fn-conflict">
          <p>The authors declared no potential conflicts of interest regarding this article’s research, authorship, and/or publication.</p>
        </fn>
      </author-notes>
      <pub-date pub-type="epub">
        <day>11</day>
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <pub-date pub-type="collection">
        <month>12</month>
        <year>2025</year>
      </pub-date>
      <volume>14</volume>
      <issue>01</issue>
      <fpage>49</fpage>
      <lpage>62</lpage>
      <history>
        <date date-type="received">
          <day>15</day>
          <month>11</month>
          <year>2025</year>
        </date>
        <date date-type="accepted">
          <day>20</day>
          <month>12</month>
          <year>2025</year>
        </date>
        <date date-type="published">
          <day>23</day>
          <month>12</month>
          <year>2025</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>© 2026 by the authors and Scientific Research Publishing Inc.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p> This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ( <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link> ). </license-p>
        </license>
      </permissions>
      <self-uri content-type="doi" xlink:href="https://doi.org/10.4236/jdaip.2026.141004">https://doi.org/10.4236/jdaip.2026.141004</self-uri>
      <abstract>
        <p>This study investigates the application of cumulative link models with alternative distributions (hyperbolic secant, Laplace, and Cauchy) to model ordinal outcomes of depressive severity using 2022 National Health Interview Survey data. The primary objective was to assess whether these models provide a better fit to ordinal response data and more accurate predictions than their traditional counterparts with the logit link function. The results indicate that the logit model achieved the highest classification accuracy, correctly classifying 83.54% of the cases. The Cauchy model demonstrated the best model fit, <italic>i.e</italic>., the lowest AIC and BIC values. This study highlights the importance of considering both classification accuracy and model fit when selecting a statistical model.</p>
      </abstract>
      <kwd-group kwd-group-type="author-generated" xml:lang="en">
        <kwd>Cumulative Link Model</kwd>
        <kwd>Ordinal Outcome</kwd>
        <kwd>Depressive Symptom Severity</kwd>
        <kwd>National Health Interview Survey</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1">
      <title>1. Introduction</title>
      <p>Ordinal response data refers to data with a categorical outcome having natural, ordered categories, but with unknown distances between these categories. Cumulative link models (CLMs) are statistical models specifically designed to analyze this data type. These models have found widespread applications in various fields. In the social sciences, for example, CLMs can be used to model attitude responses on a Likert scale [<xref ref-type="bibr" rid="B1">1</xref>]. In medicine, they can be used to analyze disease severity with cancer stage. In ecology, CLMs can be used to study the effects of factors on the spatial abundance of a species, which is categorized into distinct levels [<xref ref-type="bibr" rid="B2">2</xref>]. CLMs provide a flexible framework for modeling the relationship between an ordinal outcome and a set of predictor variables while preserving the inherent ordering of the response categories [<xref ref-type="bibr" rid="B3">3</xref>].</p>
      <p>CLMs link the cumulative probabilities of the ordinal response to a set of predictors through a suitable link function. This process assumes that the observed ordinal response is a manifestation of an underlying continuous variable with a corresponding distribution that is not directly observed. Commonly used link functions include the logit, probit, cumulative log-log, and log-log, which correspond to the logistic, normal, Gompertz, and Gumbel distributions [<xref ref-type="bibr" rid="B4">4</xref>]. The CLM with a logit link, also known as the proportional odds model, is widely used in research. These link functions impose certain assumptions on the underlying latent variable that generates the observed ordinal response. However, the choice of link function and the associated distributional assumptions can significantly influence the model’s performance and the interpretation of its results [<xref ref-type="bibr" rid="B5">5</xref>]. Extensions of CLMs include the incorporation of dispersion effects, in which the explanatory variables affect not only the location of the ordinal response but also its spread or variability [<xref ref-type="bibr" rid="B6">6</xref>].</p>
      <p>While traditional CLMs with logit or probit links are widely used, they may not always be appropriate. Each CLM assumes a distribution for the unobserved latent variable. However, this assumption may not hold when the underlying data exhibit skewness, kurtosis, or boundary inflation, leading to poorer model performance. A limitation is their inability to adequately capture complex relationships in certain situations. For instance, in surveys assessing mental health symptom severity, responses might be heavily skewed towards lower symptomatology levels [<xref ref-type="bibr" rid="B7">7</xref>]. In behavioral health studies, many patients may fall into the ‘minimal depressive symptomatology’ category. In these cases, alternative distributions potentially offering greater flexibility in modeling the shape of the latent variable distribution may be more suitable. Additional CLMs with associated distributions must be evaluated as they may better fit the data when applying the ordinal outcome model.</p>
      <p>This manuscript proposes using the hyperbolic secant, Laplace, and Cauchy distributions as a group of candidate distributions for CLMs. Integrating hyperbolic secant and Laplace distributions into CLMs represents a novel consideration within behavioral research. The hyperbolic secant distribution and its generalizations are extensively used in financial modeling; characterized by slightly fatter tails than the normal distribution, it is particularly adept at accommodating datasets with larger-than-average observations [<xref ref-type="bibr" rid="B8">8</xref>][<xref ref-type="bibr" rid="B9">9</xref>]. Moreover, this distribution has consistently demonstrated a robust fit across the entire range of data support. The Laplace distribution has also been employed in financial modeling because it captures the leptokurtic and skewed nature of financial data [<xref ref-type="bibr" rid="B10">10</xref>]. The Laplace distribution has a sharper peak at the mean than the normal distribution, but heavier tails due to slower decay. The Cauchy distribution is also bell-shaped and symmetric, but with much heavier tails than the normal distribution [<xref ref-type="bibr" rid="B11">11</xref>].</p>
      <p>Given the limitations of traditional CLMs and the flexibility offered by the distributions under consideration, this research investigates whether CLMs with hyperbolic secant, Laplace, and Cauchy distributions provide a better fit to ordinal response data. Choosing an appropriate latent distribution enables the model to capture the underlying structure of the ordinal data, leading to more accurate and consistent predictions and better model fit. Specifically, we aim to determine whether these models offer improved accuracy, model fit, and variable selection compared to their traditional counterparts with the logit link function. To achieve this, we address the following research question: </p>
      <p>Do CLMs with the hyperbolic secant, Laplace, and Cauchy distributions provide a better fit to ordinal response data and more accurate predictions than traditional CLMs with a logit link?</p>
      <p>We conducted a study and analyzed the 2022 National Health Interview Survey (NHIS) dataset with depressive symptom severity (minimal, mild, moderate, and severe) as the ordinal outcome of interest [<xref ref-type="bibr" rid="B12">12</xref>] to address the research question. The NHIS is a comprehensive, nationwide survey conducted annually by the Centers for Disease Control and Prevention that collects data on a wide range of health topics, including chronic conditions, health insurance coverage, and access to healthcare services [<xref ref-type="bibr" rid="B13">13</xref>]. The models were evaluated based on predictive accuracy, model fit, macro-average F1 score, and mean decrease in variable accuracy. The findings of this research will contribute to the development of additional CLMs that can better capture the complexities of ordinal response data.</p>
    </sec>
    <sec id="sec2">
      <title>2. Materials and Methods</title>
      <p>For a given observation <italic>i</italic>, denote <inline-formula><mml:math><mml:mrow><mml:msub><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> x </mml:mi></mml:mstyle><mml:mi> i </mml:mi></mml:msub><mml:mo> = </mml:mo><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 1 </mml:mn></mml:mrow></mml:msub><mml:mtext>   </mml:mtext><mml:mtext>   </mml:mtext><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 2 </mml:mn></mml:mrow></mml:msub><mml:mtext>   </mml:mtext><mml:mtext>   </mml:mtext><mml:mo> ⋯ </mml:mo><mml:mtext>   </mml:mtext><mml:mtext>   </mml:mtext><mml:msub><mml:mi> x </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> p </mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> as a vector of <italic>p</italic>covariates, with <inline-formula><mml:math display="inline"><mml:mrow><mml:mi> i </mml:mi><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mo> , </mml:mo><mml:mn> 2 </mml:mn><mml:mo> , </mml:mo><mml:mo> ⋯ </mml:mo><mml:mo> , </mml:mo><mml:mo> , </mml:mo><mml:mi> n </mml:mi></mml:mrow></mml:math></inline-formula> . For this study, the predictor variables comprised sociodemographic and healthcare access-related variables from the 2022 NHIS dataset. A measure of anxiety was also included as a covariate. In addition, an ordinal outcome of depressive severity was recorded. As such, we denote the outcome vector <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mo></mml:mo></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 1 </mml:mn></mml:mrow></mml:msub><mml:mo> , </mml:mo><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 2 </mml:mn></mml:mrow></mml:msub><mml:mo> , </mml:mo><mml:mo> ⋯ </mml:mo><mml:mo> , </mml:mo><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> J </mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:math></inline-formula> if, for observation <italic>i</italic>, the outcome is in the <italic>j</italic><sup>th</sup> category, with all other entries being set to 0. There are J possible outcome levels, with depressive severity measured by the PHQ-8. These levels are categorized as follows:</p>
      <p>1. Minimal (PHQ-8 score under five)</p>
      <p>2. Mild (PHQ-8 score from five to less than 10)</p>
      <p>3. Moderate (PHQ-8 score from 10 to less than 15)</p>
      <p>4. Severe (PHQ-8 score of 15 or higher)</p>
      <p>For this given study, the goal is to use the covariate vector <inline-formula><mml:math><mml:mrow><mml:msub><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> x </mml:mi></mml:mstyle><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> to predict depressive severity, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> y </mml:mi></mml:mstyle><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> , using the 2022 NHIS data. The aggregation of vectors for all subjects yields the covariate matrix <inline-formula><mml:math><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> X </mml:mi></mml:mstyle></mml:math></inline-formula> , a<italic>p</italic>by <italic>n</italic> matrix where the <italic>i</italic><italic><sup>th</sup></italic> column is set to <inline-formula><mml:math><mml:mrow><mml:msub><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> x </mml:mi></mml:mstyle><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> . We also have a <italic>J</italic>by <italic>n</italic>matrix, denoted <inline-formula><mml:math><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> Y </mml:mi></mml:mstyle></mml:math></inline-formula> , where the <italic>i</italic><italic><sup>th</sup></italic> column is set to <inline-formula><mml:math><mml:mrow><mml:msub><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> y </mml:mi></mml:mstyle><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> . The goal is to develop a function <inline-formula><mml:math><mml:mrow><mml:mi> f </mml:mi><mml:mo> : </mml:mo><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> X </mml:mi></mml:mstyle><mml:mo> → </mml:mo><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> Y </mml:mi></mml:mstyle></mml:mrow></mml:math></inline-formula> , which aims to predict the <inline-formula><mml:math><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> Y </mml:mi></mml:mstyle></mml:math></inline-formula> matrix using the <inline-formula><mml:math><mml:mstyle mathvariant="bold" mathsize="normal"><mml:mi> X </mml:mi></mml:mstyle></mml:math></inline-formula> matrix. To achieve this, an ordinal regression framework was used. First, four CLMs are introduced, two of which are novel applications (based on the hyperbolic secant and Laplace distributions). Next, the models’ specifications are outlined. The method is then applied to the 2022 NHIS data, with results on predictive accuracy, macro-averaged F1 score, model fit, and variable importance reported.</p>
      <sec id="sec2dot1">
        <title>2.1. Cumulative Link-Based Outcome Functions</title>
        <p>The cost function for the neural network is derived from the log-likelihood of a multinomial distribution:</p>
        <disp-formula id="FD1">
          <label>(1)</label>
          <mml:math>
            <mml:mrow>
              <mml:mi>log</mml:mi>
              <mml:mi>L</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mstyle displaystyle="true">
                <mml:munderover>
                  <mml:mo>∑</mml:mo>
                  <mml:mrow>
                    <mml:mi>i</mml:mi>
                    <mml:mo>=</mml:mo>
                    <mml:mn>1</mml:mn>
                  </mml:mrow>
                  <mml:mi>n</mml:mi>
                </mml:munderover>
                <mml:mrow>
                  <mml:mstyle displaystyle="true">
                    <mml:munderover>
                      <mml:mo>∑</mml:mo>
                      <mml:mrow>
                        <mml:mi>j</mml:mi>
                        <mml:mo>=</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                      <mml:mi>J</mml:mi>
                    </mml:munderover>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>y</mml:mi>
                        <mml:mrow>
                          <mml:mi>i</mml:mi>
                          <mml:mi>j</mml:mi>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>×</mml:mo>
                      <mml:mi>log</mml:mi>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>π</mml:mi>
                            <mml:mrow>
                              <mml:mi>i</mml:mi>
                              <mml:mi>j</mml:mi>
                            </mml:mrow>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:mstyle>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> π </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mi> P </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mo> | </mml:mo><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> . Define the cumulative probabilities <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> as </p>
        <p><inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mi> P </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mo> | </mml:mo><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> + </mml:mo><mml:mi> P </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi><mml:mo> − </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mo> | </mml:mo><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow><mml:mo> + </mml:mo><mml:mo> ⋯ </mml:mo><mml:mo> + </mml:mo><mml:mi> P </mml:mi><mml:mrow><mml:mo> ( </mml:mo><mml:mrow><mml:msub><mml:mi> y </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 1 </mml:mn></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn><mml:mo> | </mml:mo><mml:msub><mml:mi> x </mml:mi><mml:mi> i </mml:mi></mml:msub></mml:mrow><mml:mo> ) </mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> . In addition, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> J </mml:mi></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 1 </mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math display="inline"><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mn> 0 </mml:mn></mml:mrow></mml:msub><mml:mo> = </mml:mo><mml:mn> 0 </mml:mn></mml:mrow></mml:math></inline-formula> . As such, we can write Equation (1) as </p>
        <disp-formula id="FD2">
          <label>(2)</label>
          <mml:math>
            <mml:mrow>
              <mml:mi>log</mml:mi>
              <mml:mi>L</mml:mi>
              <mml:mo>=</mml:mo>
              <mml:mstyle displaystyle="true">
                <mml:munderover>
                  <mml:mo>∑</mml:mo>
                  <mml:mrow>
                    <mml:mi>i</mml:mi>
                    <mml:mo>=</mml:mo>
                    <mml:mn>1</mml:mn>
                  </mml:mrow>
                  <mml:mi>n</mml:mi>
                </mml:munderover>
                <mml:mrow>
                  <mml:mstyle displaystyle="true">
                    <mml:munderover>
                      <mml:mo>∑</mml:mo>
                      <mml:mrow>
                        <mml:mi>j</mml:mi>
                        <mml:mo>=</mml:mo>
                        <mml:mn>1</mml:mn>
                      </mml:mrow>
                      <mml:mi>J</mml:mi>
                    </mml:munderover>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>y</mml:mi>
                        <mml:mrow>
                          <mml:mi>i</mml:mi>
                          <mml:mi>j</mml:mi>
                        </mml:mrow>
                      </mml:msub>
                      <mml:mo>×</mml:mo>
                      <mml:mi>log</mml:mi>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>p</mml:mi>
                            <mml:mrow>
                              <mml:mi>i</mml:mi>
                              <mml:mi>j</mml:mi>
                            </mml:mrow>
                          </mml:msub>
                          <mml:mo>−</mml:mo>
                          <mml:msub>
                            <mml:mi>p</mml:mi>
                            <mml:mrow>
                              <mml:mi>i</mml:mi>
                              <mml:mi>j</mml:mi>
                              <mml:mo>−</mml:mo>
                              <mml:mn>1</mml:mn>
                            </mml:mrow>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:mstyle>
                </mml:mrow>
              </mml:mstyle>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>We are primarily concerned with directly modeling cumulative probabilities and, to this end, employed CLMs.</p>
        <p>CLMs are statistical models intended to analyze data that have ordinal outcomes. We are concerned with modeling the cumulative probabilities <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> . In CLMs, <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> p </mml:mi><mml:mrow><mml:mi> i </mml:mi><mml:mi> j </mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is linked to a function of predictor variables through a specified link function. This method enables the use of a covariate set to account for variation in the ordinal outcome while preserving the natural order of the response categories. The four-link functions considered in this study are based on the following distributions:</p>
        <p>1. Logistic Distribution (logit)</p>
        <p>2. Hyperbolic Secant</p>
        <p>3. Cauchy</p>
        <p>4. Laplace</p>
        <p>These special distributions (hyperbolic secant, Laplace, and Cauchy), rather than the more common alternatives (normal (probit) and Gumbel (complementary log-log)), were chosen due to their tail and kurtosis properties to allow greater flexibility to accommodate the skewed and heavy-tailed characteristics of the behavioral health data that the other distributions may not capture. </p>
        <p>The goal of this study is to evaluate the performance of the link functions with respect to predictive accuracy (defined as the model’s performance on a validation dataset), model fit, and variable selection (defined as the mean decrease in accuracy).</p>
        <p>By employing the logit link, the cumulative probabilities are modeled as:</p>
        <disp-formula id="FD3">
          <label>(3)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msubsup>
                <mml:mi>p</mml:mi>
                <mml:mrow>
                  <mml:mi>i</mml:mi>
                  <mml:mi>j</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>[</mml:mo>
                    <mml:mn>1</mml:mn>
                    <mml:mo>]</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:msubsup>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mrow>
                  <mml:mn>1</mml:mn>
                  <mml:mo>+</mml:mo>
                  <mml:msup>
                    <mml:mtext>e</mml:mtext>
                    <mml:mrow>
                      <mml:mo>−</mml:mo>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mstyle mathvariant="bold" mathsize="normal">
                              <mml:mi>x</mml:mi>
                            </mml:mstyle>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                          <mml:msub>
                            <mml:mi>β</mml:mi>
                            <mml:mi>j</mml:mi>
                          </mml:msub>
                          <mml:mo>+</mml:mo>
                          <mml:msub>
                            <mml:mi>b</mml:mi>
                            <mml:mi>j</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:msup>
                </mml:mrow>
              </mml:mfrac>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>where <inline-formula><mml:math><mml:mrow><mml:msub><mml:mi> b </mml:mi><mml:mi> j </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is defined as the intercept parameter for the <italic>j</italic>th level. Considering the hyperbolic secant distribution as the latent distribution of interest, the cumulative probabilities are now modeled as:</p>
        <disp-formula id="FD4">
          <label>(4)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msubsup>
                <mml:mi>p</mml:mi>
                <mml:mrow>
                  <mml:mi>i</mml:mi>
                  <mml:mi>j</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>[</mml:mo>
                    <mml:mn>2</mml:mn>
                    <mml:mo>]</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:msubsup>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mn>2</mml:mn>
                <mml:mi>π</mml:mi>
              </mml:mfrac>
              <mml:mi>arctan</mml:mi>
              <mml:mrow>
                <mml:mo>[</mml:mo>
                <mml:mrow>
                  <mml:mi>exp</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mfrac>
                        <mml:mi>π</mml:mi>
                        <mml:mn>2</mml:mn>
                      </mml:mfrac>
                      <mml:mrow>
                        <mml:mo>{</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>x</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                          <mml:msub>
                            <mml:mi>β</mml:mi>
                            <mml:mi>j</mml:mi>
                          </mml:msub>
                          <mml:mo>+</mml:mo>
                          <mml:msub>
                            <mml:mi>b</mml:mi>
                            <mml:mi>j</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>}</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mrow>
                <mml:mo>]</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>When the underlying latent distribution is assumed to be a Cauchy distribution, we have</p>
        <disp-formula id="FD5">
          <label>(5)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:msubsup>
                <mml:mi>p</mml:mi>
                <mml:mrow>
                  <mml:mi>i</mml:mi>
                  <mml:mi>j</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:mrow>
                    <mml:mo>[</mml:mo>
                    <mml:mn>3</mml:mn>
                    <mml:mo>]</mml:mo>
                  </mml:mrow>
                </mml:mrow>
              </mml:msubsup>
              <mml:mo>=</mml:mo>
              <mml:mfrac>
                <mml:mn>1</mml:mn>
                <mml:mi>π</mml:mi>
              </mml:mfrac>
              <mml:mi>arctan</mml:mi>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>x</mml:mi>
                    <mml:mi>i</mml:mi>
                  </mml:msub>
                  <mml:msub>
                    <mml:mi>β</mml:mi>
                    <mml:mi>j</mml:mi>
                  </mml:msub>
                  <mml:mo>+</mml:mo>
                  <mml:msub>
                    <mml:mi>b</mml:mi>
                    <mml:mi>j</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>+</mml:mo>
              <mml:mn>0.5</mml:mn>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>Finally, utilizing the Laplace distribution as the underlying latent distribution, we have:</p>
        <disp-formula id="FD6">
          <label>(6)</label>
          <mml:math display="inline">
            <mml:mtable>
              <mml:mtr>
                <mml:mtd>
                  <mml:msubsup>
                    <mml:mi>p</mml:mi>
                    <mml:mrow>
                      <mml:mi>i</mml:mi>
                      <mml:mi>j</mml:mi>
                    </mml:mrow>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>[</mml:mo>
                        <mml:mn>4</mml:mn>
                        <mml:mo>]</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                  </mml:msubsup>
                  <mml:mo>=</mml:mo>
                  <mml:mi>I</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>x</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:msub>
                        <mml:mi>β</mml:mi>
                        <mml:mi>j</mml:mi>
                      </mml:msub>
                      <mml:mo>+</mml:mo>
                      <mml:msub>
                        <mml:mi>b</mml:mi>
                        <mml:mi>j</mml:mi>
                      </mml:msub>
                      <mml:mo>≥</mml:mo>
                      <mml:mn>0</mml:mn>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                </mml:mtd>
              </mml:mtr>
              <mml:mtr>
                <mml:mtd>
                  <mml:mo>−</mml:mo>
                  <mml:mi>s</mml:mi>
                  <mml:mi>i</mml:mi>
                  <mml:mi>g</mml:mi>
                  <mml:mi>n</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:msub>
                        <mml:mi>x</mml:mi>
                        <mml:mi>i</mml:mi>
                      </mml:msub>
                      <mml:msub>
                        <mml:mi>β</mml:mi>
                        <mml:mi>j</mml:mi>
                      </mml:msub>
                      <mml:mo>+</mml:mo>
                      <mml:msub>
                        <mml:mi>b</mml:mi>
                        <mml:mi>j</mml:mi>
                      </mml:msub>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mfrac>
                    <mml:mn>1</mml:mn>
                    <mml:mn>2</mml:mn>
                  </mml:mfrac>
                  <mml:mi>exp</mml:mi>
                  <mml:mrow>
                    <mml:mo>(</mml:mo>
                    <mml:mrow>
                      <mml:mrow>
                        <mml:mo>{</mml:mo>
                        <mml:mrow>
                          <mml:msub>
                            <mml:mi>x</mml:mi>
                            <mml:mi>i</mml:mi>
                          </mml:msub>
                          <mml:msub>
                            <mml:mi>β</mml:mi>
                            <mml:mi>j</mml:mi>
                          </mml:msub>
                          <mml:mo>+</mml:mo>
                          <mml:msub>
                            <mml:mi>b</mml:mi>
                            <mml:mi>j</mml:mi>
                          </mml:msub>
                        </mml:mrow>
                        <mml:mo>}</mml:mo>
                      </mml:mrow>
                      <mml:mo>×</mml:mo>
                      <mml:mrow>
                        <mml:mo>(</mml:mo>
                        <mml:mrow>
                          <mml:mo>−</mml:mo>
                          <mml:mi>s</mml:mi>
                          <mml:mi>i</mml:mi>
                          <mml:mi>g</mml:mi>
                          <mml:mi>n</mml:mi>
                          <mml:mrow>
                            <mml:mo>(</mml:mo>
                            <mml:mrow>
                              <mml:msub>
                                <mml:mi>x</mml:mi>
                                <mml:mi>i</mml:mi>
                              </mml:msub>
                              <mml:msub>
                                <mml:mi>β</mml:mi>
                                <mml:mi>j</mml:mi>
                              </mml:msub>
                              <mml:mo>+</mml:mo>
                              <mml:msub>
                                <mml:mi>b</mml:mi>
                                <mml:mi>j</mml:mi>
                              </mml:msub>
                            </mml:mrow>
                            <mml:mo>)</mml:mo>
                          </mml:mrow>
                        </mml:mrow>
                        <mml:mo>)</mml:mo>
                      </mml:mrow>
                    </mml:mrow>
                    <mml:mo>)</mml:mo>
                  </mml:mrow>
                  <mml:mo>,</mml:mo>
                </mml:mtd>
              </mml:mtr>
            </mml:mtable>
          </mml:math>
        </disp-formula>
        <p>These cumulative probabilities can be substituted into equation (2) and solved accordingly. The goal is to find parameter estimates such that:</p>
        <disp-formula id="FD7">
          <label>(7)</label>
          <mml:math display="inline">
            <mml:mrow>
              <mml:mrow>
                <mml:mo>(</mml:mo>
                <mml:mrow>
                  <mml:msub>
                    <mml:mover accent="true">
                      <mml:mi>β</mml:mi>
                      <mml:mo>^</mml:mo>
                    </mml:mover>
                    <mml:mi>j</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mover accent="true">
                      <mml:mi>b</mml:mi>
                      <mml:mo>^</mml:mo>
                    </mml:mover>
                    <mml:mi>j</mml:mi>
                  </mml:msub>
                </mml:mrow>
                <mml:mo>)</mml:mo>
              </mml:mrow>
              <mml:mo>=</mml:mo>
              <mml:munder>
                <mml:mrow>
                  <mml:mi>arg</mml:mi>
                  <mml:mi>min</mml:mi>
                </mml:mrow>
                <mml:mrow>
                  <mml:msub>
                    <mml:mi>β</mml:mi>
                    <mml:mi>j</mml:mi>
                  </mml:msub>
                  <mml:mo>,</mml:mo>
                  <mml:msub>
                    <mml:mi>b</mml:mi>
                    <mml:mi>j</mml:mi>
                  </mml:msub>
                </mml:mrow>
              </mml:munder>
              <mml:mrow>
                <mml:mo>{</mml:mo>
                <mml:mrow>
                  <mml:mo>−</mml:mo>
                  <mml:mi>log</mml:mi>
                  <mml:mi>L</mml:mi>
                </mml:mrow>
                <mml:mo>}</mml:mo>
              </mml:mrow>
            </mml:mrow>
          </mml:math>
        </disp-formula>
        <p>As such, the Adam optimization algorithm was applied [<xref ref-type="bibr" rid="B14">14</xref>][<xref ref-type="bibr" rid="B15">15</xref>]. Once the optimal values for <inline-formula><mml:math display="inline"><mml:mover accent="true"><mml:mi> β </mml:mi><mml:mo> ^ </mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> b </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mi> j </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> were computed, they were used to evaluate the models regarding accuracy, model fit, and variable selection.</p>
      </sec>
      <sec id="sec2dot2">
        <title>2.2. Application to NHIS Data</title>
        <p>The NHIS is a pivotal, cross-sectional household survey conducted annually by the National Center for Health Statistics (NCHS) under the Centers for Disease Control and Prevention (CDC) [<xref ref-type="bibr" rid="B16">16</xref>]. The primary aim of this survey is to monitor the health status of the U.S. population and to track trends in essential health indicators. The survey gathers comprehensive data on a broad spectrum of health-related topics, including chronic and acute conditions, health insurance coverage, utilization of healthcare services, and health-related behaviors. Given that the NHIS provides a nationally representative sample of the civilian noninstitutionalized population, its data are widely used by researchers and policymakers to assess public health needs, evaluate health policies, and inform interventions to enhance Americans’ health [<xref ref-type="bibr" rid="B16">16</xref>]. In the context of modeling depressive severity, the NHIS is particularly valuable due to its use of the PHQ-8, a widely accepted screening tool for depressive symptoms [<xref ref-type="bibr" rid="B17">17</xref>]. The instrument is a truncated version of the PHQ-9, aligns with DSM-IV criteria [<xref ref-type="bibr" rid="B18">18</xref>], and assesses symptom frequency over the past two weeks, enabling the determination of depression severity [<xref ref-type="bibr" rid="B19">19</xref>]. The survey’s rich collection of sociodemographic and health-related variables further enables researchers to explore how a variety of factors, from income and education to access to care, are associated with depression severity across a nationally representative sample of the population.</p>
        <p>For this study, the 2022 survey data were used. The outcome variable is an ordinal measure of the PHQ-8, as described earlier in the methods section. Due to the relatively small sample sizes at moderate and severe levels, these two levels were combined into a single moderate/severe level. Selection of specific predictor variables was guided by established literature on sociodemographic and healthcare access-related risk factors for depression [<xref ref-type="bibr" rid="B20">20</xref>]. The predictor variables include age, poverty levels (less than 100% Federal Poverty Level (FPL), between 100% and 199% FPL, between 200% and 300% FPL, and greater than 400% FPL), sex (male, female), race/ethnicity (non-Hispanic white, non-Hispanic black, other), education level (some high school, high school graduate, some college, bachelor’s, master’s, professional or doctoral degree), health insurance (private, Medicare, other, uninsured), any delay in receiving medical care over the past 12 months (yes, no), foregoing medical care due to cost in the past 12 months (yes, no), having a usual place for medical care (yes, no), the number of urgent care visits in the past year (0, 1, &gt;1), the number of emergency room visits in the past year (0, 1, &gt;1), any overnight hospitalizations in the past year (yes, no), living alone (yes, no), and GAD-7. The GAD-7 is represented ordinally as:</p>
        <p>1. Minimal (GAD-7 score less than five)</p>
        <p>2. Mild (GAD-7 score greater than or equal to five and less than 10)</p>
        <p>3. Moderate (GAD-7 score greater than or equal to 10 and less than 15)</p>
        <p>4. Severe (GAD-7 score greater than or equal to 15)</p>
        <p>To evaluate the models on the 2022 NHIS dataset, the data was split into a training set (80%) and a validation set (20%). The model parameters for the four cumulative link regression models were optimized on the training dataset. Once optimized, the model parameters <inline-formula><mml:math><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> β </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mi> j </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi> b </mml:mi><mml:mo> ^ </mml:mo></mml:mover><mml:mi> i </mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> were applied and evaluated on the validation data. The metrics reported include the accuracy (percent correctly classified and percent correctly classified in the moderate/severe depressive symptomatology category), Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), and macro-averaged F1 score [<xref ref-type="bibr" rid="B21">21</xref>], and the mean decrease in accuracy for each variable. The macro-average F1 score computes the F1 score for each class in a multi-class classification problem and then averages those scores. It gives equal weight to all classes, regardless of the number of samples in each class, and is useful for comparing performance on imbalanced datasets where some classes may have few examples. The mean decrease in accuracy (also known as Permutation Feature Importance) is a technique for quantifying the importance of a feature by measuring the average drop in a model’s prediction accuracy when the feature values are randomly permuted [<xref ref-type="bibr" rid="B22">22</xref>]. Permuting the feature values breaks the relationship between the feature values and the true outcome, effectively excluding it from the model. This process was repeated 100 times for each variable, with the average measure being reported. A high score indicates that the model is highly reliant on that feature, as its performance drops significantly when the feature values are randomly permuted. All analyses were performed using the R software program [<xref ref-type="bibr" rid="B23">23</xref>].</p>
      </sec>
    </sec>
    <sec id="sec3">
      <title>3. Results</title>
      <p>The study applied four cumulative link models to the 2022 NHIS dataset to evaluate their performance in modeling depressive severity. The sample size for this study is 25,208. The models included logit, hyperbolic secant, Cauchy, and Laplace distributions. The primary evaluation metrics were the percentage of correctly classified cases with 95% confidence intervals, AIC, BIC, macro-average F1 score, and variable importance measured by the decrease in accuracy. <bold>Table 1</bold> illustrates the distribution of the ordinal outcome within the National Health Interview Survey (NHIS) dataset. A substantial proportion of participants exhibited minimal depressive symptomatology, accounting for 79.07% of the sample. In contrast, a smaller fraction of individuals reported moderate and severe symptomatology, comprising 4.36% and 2.8% of the sample, respectively. Consequently, for the purposes of further modeling analysis, the moderate and severe categories were amalgamated into a single category.</p>
      <p><bold>Table 1.</bold> Descriptive statistics for the ordinal outcome of PHQ-8 symptom severity.</p>
      <table-wrap id="tbl1">
        <label>Table 1</label>
        <table>
          <tbody>
            <tr>
              <td>
              </td>
              <td colspan="4">PHQ-8 (Categorized)</td>
            </tr>
            <tr>
              <td>
              </td>
              <td>Minimal</td>
              <td>Mild</td>
              <td>Moderate</td>
              <td>Severe</td>
            </tr>
            <tr>
              <td>N (%)</td>
              <td>19933 (79.07%)</td>
              <td>3468 (13.76%)</td>
              <td>1100 (4.36%)</td>
              <td>707 (2.8%)</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <p>Due to the relatively small sample size, moderate and severe were combined into a single group.</p>
      <p><bold>Table 2.</bold> Ordinal regression network accuracy.</p>
      <table-wrap id="tbl2">
        <label>Table 2</label>
        <table>
          <tbody>
            <tr>
              <td>
              </td>
              <td colspan="4">
                <bold>Model Type</bold>
              </td>
            </tr>
            <tr>
              <td>
              </td>
              <td>Logit</td>
              <td>Hyperbolic Secant</td>
              <td>Cauchy</td>
              <td>Laplace</td>
            </tr>
            <tr>
              <td>
                <bold>Correctly Classified (95% Confidence Interval)</bold>
              </td>
              <td>83.54% (82.49%, 84.55%)</td>
              <td>83.42% (82.36%, 84.44%)</td>
              <td>80.78% (79.67%, 81.86%)</td>
              <td>83.16% (82.1%, 84.18%)</td>
            </tr>
            <tr>
              <td>
                <bold>AIC</bold>
              </td>
              <td>7511.46</td>
              <td>11830.76</td>
              <td>6127.67</td>
              <td>8722.97</td>
            </tr>
            <tr>
              <td>
                <bold>BIC</bold>
              </td>
              <td>7655.4</td>
              <td>11974.7</td>
              <td>6271.61</td>
              <td>8866.91</td>
            </tr>
            <tr>
              <td>
                <bold>Macro-averaged F1 Score</bold>
              </td>
              <td>0.62</td>
              <td>0.61</td>
              <td>0.49</td>
              <td>0.6</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <p>The percentage correctly classified by the four models (logit, hyperbolic secant, Cauchy, and Laplace) applied to the 2022 National Health Interview Survey data.</p>
      <p><bold>Table 2</bold> reports the metrics regarding predictive accuracy and model fit as applied to the validation dataset. Regarding model performance in terms of correctly classified percentage, the logit model achieved the highest accuracy, correctly classifying 83.54% of the cases studied. Closely following, the hyperbolic secant and Laplace models achieved accuracies of 83.42% and 83.16%, respectively. Conversely, the Cauchy model exhibited the lowest accuracy, correctly classifying 80.78% of the cases. Lower AIC and BIC scores indicate better model performance. Interestingly, the Cauchy model performed best, presenting the lowest AIC (6127.67) and BIC (6271.61) values, surpassing the other models. In contrast, the logit model yielded AIC and BIC values of 7511.46 and 7655.40, respectively, while the hyperbolic secant and Laplace models recorded even higher AIC and BIC values, indicating poorer performance than the Cauchy model. Regarding the macro-averaged F1 score, the logit model achieved the highest score, followed by the hyperbolic secant and Laplace-based models; the Cauchy model had the lowest score.</p>
      <p>To assess the models’ predictive capability at the extremes, predictive accuracy was evaluated in the minority class of moderate/severe depressive symptom severity. The percent correctly classified on the validation data is listed as follows:</p>
      <p>1. Logistic: 53.37%</p>
      <p>2. Hyperbolic Secant: 56.13%</p>
      <p>3. Cauchy: 17.79%</p>
      <p>4. Laplace: 53.06%</p>
      <p>As such, the hyperbolic secant outperformed all other models when predicting the extreme class of moderate/severe depression symptom severity.</p>
      <fig id="fig1">
        <label>Figure 1</label>
        <graphic xlink:href="https://html.scirp.org/file/2870863-rId77.jpeg?20251223020419" />
      </fig>
      <p><bold>Figure 1.</bold> The plot presents the mean decrease in accuracy per variable for the four models presented. The x-axis represents decrease in accuracy, while the y-axis displays the variables used in the model.</p>
      <p><xref ref-type="fig" rid="fig1">Figure 1</xref> shows the average decrease in accuracy across the four models tested on the validation dataset. The mean decrease in accuracy was used to evaluate the importance of predictor variables in each model. When analyzing the mean decrease in accuracy across the four models, the GAD-7 severity variable consistently shows the largest decrease. This indicates that GAD-7 has the greatest effect on model performance. In both the Laplace and hyperbolic secant models, a similar pattern emerges: variables such as past 12-month delay in receiving medical care and past-year number of emergency room visits are associated with small drops in accuracy. Conversely, within the Cauchy model framework, the variables health insurance and age have a more significant influence than in the other models, though still minor. The logit model shows that the variables—any overnight hospitalizations in the past year and any delay in receiving medical care in the past 12 months—are the next most important factors affecting accuracy, after GAD-7.</p>
    </sec>
    <sec id="sec4">
      <title>4. Discussion</title>
      <p>This study explored the application of CLMs with alternative distributions—hyperbolic secant, Laplace, and Cauchy—to model ordinal outcomes of depressive severity using 2022 NHIS data. The primary objective was to assess whether these models provide a better fit and more accurate predictions than traditional CLMs with a logit link.</p>
      <p>The results indicate that the logit model achieved the highest classification accuracy. The logit also achieved the highest macro-averaged F1 score. The hyperbolic secant-based model achieved the highest predictive accuracy when predicting the minority class of moderate/severe depression symptom severity. A plausible reason is that the logistic cumulative distribution function naturally arises when the log-likelihood of the multinomial distribution is derived. Also, the logit link function performed well in estimating the threshold for the minimal category. However, the Cauchy model demonstrated superior model fit, as evidenced by the lowest AIC and BIC values. A possible hypothesis for the improved performance of the Cauchy’s heavy tails is that they do a better job of accommodating the individual whose depression symptom severity scores are in the moderate/severe range. The tails of the Cauchy distribution decay more slowly than those of the logistic distribution, so it might be able to model the extreme observation of moderate/severe depression symptom severity more accurately without being overly influenced by them, leading to a higher likelihood value, and lower AIC and BIC across the dataset. The Cauchy distribution may sacrifice accuracy, but it gains by better representing the true distributional shape, especially at the extremes. These findings suggest that while the logit model is effective for classification tasks, the Cauchy model may offer a more nuanced understanding of the data structure, particularly when model fit is prioritized.</p>
      <p>The trade-off between the highest predictive accuracy and better model fit can have significant practical implications for health researchers: choices about which model to use should depend on the study’s main goal. For instance, in a healthcare setting, if the objective is to develop a tool that accurately classifies patients into depression severity categories (such as for automated screening or triage), the model that correctly classifies the most cases should be preferred. However, if the goal of a study is to gain a deeper or more nuanced understanding of the relationship between the predictors and the underlying latent variable (like depression severity), model fit (AIC and BIC) should be prioritized.</p>
      <p>The variable importance analysis showed that the GAD-7 severity variable consistently had the highest mean decrease in accuracy across all models, indicating its significant impact on predicting depressive severity. The logit and hyperbolic secant models also identified any delay in receiving medical care over the past 12 months and the number of emergency room visits in the past year as key factors, while the Cauchy model emphasized the importance of health insurance and age. This difference in variable importance across models suggests that various distributions may capture different aspects of the data, providing complementary insights.</p>
      <p>We acknowledge limitations due to data preparation, namely, the requirement to combine the moderate and severe depressive symptom categories into a single moderate/severe symptom category (due to small sample sizes in the validation dataset), which assured model fitting power and stability, but potentially masks meaningful and clinically relevant differences between moderate and severe depression symptom severity. All future interpretations of the results need to be cautious about the exact clinical meaning of the moderate/severe category, because the model’s outputs and variable importance for the moderate/severe outcome reflect the merged group rather than two distinct clinical entities.</p>
      <p>The findings underscore the potential of alternative distributions, such as the Cauchy, to capture the complexities of ordinal response data, particularly in behavioral health research. The ability of these models to accommodate skewness and heavy tails makes them suitable for datasets with extreme observations, which are common in mental health surveys. Future research could explore integrating these CLMs with machine learning techniques such as penalized modeling, deep learning, and explainable AI (XAI) to further enhance predictive accuracy, model interpretability, and explainability. </p>
      <p>Despite the higher-level performance of the logit and Cauchy functions compared to the hyperbolic secant and Laplace functions, these latter functions retain their value, particularly in the context of sensitivity analyses. These analyses are crucial for validating the assumption that the underlying latent variable conforms to a specified distribution. In the health sciences, the logit link function is predominantly used for ordinal regression because it produces regression coefficients that are readily interpretable as odds ratios [<xref ref-type="bibr" rid="B24">24</xref>]. However, the other three models should be considered as viable alternatives, as it is entirely possible that they may more closely align with the underlying latent distribution. </p>
      <p>Additionally, there was a high-class imbalance in the ordinal outcome variable of depression symptom severity as measured by the PHQ-8, and some of the input variables were ordinal as well. A possible remedy to this was the application of Traditional Synthetic Minority Oversampling Technique (SMOTE) algorithms to address the class imbalance. However, for this study, SMOTE algorithms could not be directly applied, as none of the current versions accommodate ordinal outcomes and ordinal input variables [<xref ref-type="bibr" rid="B25">25</xref>]. Currently, there are no SMOTE-based algorithms in R or Python that can accommodate ordinal input and outcome variables. Additional research in algorithmic development is needed to develop SMOTE algorithms that can accommodate the complexities of high-dimensional biomedical data. A future SMOTE implementation designed to effectively address the methodological gap of class imbalance in ordinal data would require the following specific features and capabilities:</p>
      <p>1) Ordinal-Aware Distance Metric: The core capability is a distance function that respects the ordered, non-numeric nature of ordinal variables. For example, the distance between categories 1 and 2 is smaller than that between 1 and 4, but it is not calculated by simple subtraction, as with continuous variables. The algorithm must use metrics that consider cumulative probabilities or employ rank-based distances. </p>
      <p>2) Synthetic Sample Generation for Ordinal Outcomes: The algorithm must synthesize new minority-class samples for the ordinal outcome, ensuring they maintain the inherent ordering structure of the outcome variable. New synthetic outcomes should fit logically within existing categories (e.g., a new synthetic severity score should be labeled “Mild” or “Moderate”). </p>
      <p>Previous studies have demonstrated the effectiveness of the SMOTE algorithm in improving prediction accuracy for unbalanced data with binary outcomes and no ordinal input variables [<xref ref-type="bibr" rid="B26">26</xref>].</p>
      <p>This study highlights the importance of considering both classification accuracy and model fit when selecting a statistical model. Researchers should weigh these factors based on the specific objectives of their study, whether they aim to maximize predictive accuracy or to gain deeper insights into model fit. For future project implementation, one can begin by clearly defining the research question, such as whether the focus is on predictive accuracy, accuracy within a specific subgroup, or overall model fit. If pilot, preliminary, or real-world data are available, the candidate models can be evaluated, and the model with the highest performance on the prespecified metric of interest will be selected as the main model for the primary analysis of the main study. Alternative models can be used for secondary and sensitivity analyses. It is also helpful to consider additional CLMs, including probit, cloglog, and loglog. Additionally, the extensive nature of the NHIS data facilitates the examination of health from a caring science perspective, offering insights into the humanistic and relational dimensions of healthcare [<xref ref-type="bibr" rid="B27">27</xref>]. This aligns with the core concept of health as the fundamental category of caring, aiming to support and strengthen an individual’s health processes [<xref ref-type="bibr" rid="B28">28</xref>]. </p>
    </sec>
    <sec id="sec5">
      <title>5. Conclusion</title>
      <p>In conclusion, this study highlighted that using CLMs with alternative distributions, such as the hyperbolic secant, Laplace, and Cauchy, offers a promising avenue for improving the analysis of ordinal outcomes in health research. Use of such models may provide a better fit to the latent distribution than traditional models, such as the logit link, can capture effectively due to skewness and heavy tails in real-life data. In the present study, we observed that although the logit link model had the highest classification accuracy and macro-averaged F1 score, the hyperbolic secant model demonstrated the highest accuracy in the extreme cases, and the Cauchy model had the best model fit, <italic>i.e.</italic>, the lowest AIC and BIC values. This shows that model selection depends on the study’s objective: whether model fit or predictive accuracy is the primary focus. Further, variation in variable importance across models shows the potential of these alternative distributions to capture different aspects of the data, which could enrich data analysts’ options for choosing an appropriate model. By expanding the toolkit of available models, researchers can better address the complexities inherent in real-world data, ultimately leading to more informed decision-making.</p>
    </sec>
    <sec id="sec6">
      <title>Data Availability Statement</title>
      <p>The data are accessed via the URL: <ext-link ext-link-type="uri" xlink:href="https://www.cdc.gov/nchs/nhis/documentation/2022-nhis.html">https://www.cdc.gov/nchs/nhis/documentation/2022-nhis.html</ext-link>.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <title>References</title>
      <ref id="B1">
        <label>1.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Larasati, A., DeYong, C. and Slevitch, L. (2011) Comparing Neural Network and Ordinal Logistic Regression to Analyze Attitude Responses. <italic>Service</italic><italic>Science</italic>, 3, 304-312. https://doi.org/10.1287/serv.3.4.304 <pub-id pub-id-type="doi">10.1287/serv.3.4.304</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1287/serv.3.4.304">https://doi.org/10.1287/serv.3.4.304</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Larasati, A.</string-name>
              <string-name>DeYong, C.</string-name>
              <string-name>Slevitch, L.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Comparing Neural Network and Ordinal Logistic Regression to Analyze Attitude Responses</article-title>
            <source>Service Science</source>
            <volume>3</volume>
            <pub-id pub-id-type="doi">10.1287/serv.3.4.304</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B2">
        <label>2.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Guisan, A. and Harrell, F.E. (2000) Ordinal Response Regression Models in Ecology. <italic>Journal</italic><italic>of</italic><italic>Vegetation</italic><italic>Science</italic>, 11, 617-626. https://doi.org/10.2307/3236568 <pub-id pub-id-type="doi">10.2307/3236568</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2307/3236568">https://doi.org/10.2307/3236568</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Guisan, A.</string-name>
              <string-name>Harrell, F.E.</string-name>
            </person-group>
            <year>2000</year>
            <article-title>Ordinal Response Regression Models in Ecology</article-title>
            <source>Journal of Vegetation Science</source>
            <volume>11</volume>
            <pub-id pub-id-type="doi">10.2307/3236568</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B3">
        <label>3.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Christensen, R.H.B. (2018) Cumulative Link Models for Ordinal Regression with the R Package Ordinal. <italic>Journal of Statistical Software</italic>, 35, 1-46.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Christensen, R.H.B.</string-name>
            </person-group>
            <year>2018</year>
            <article-title>Cumulative Link Models for Ordinal Regression with the R Package Ordinal</article-title>
            <source>Journal of Statistical Software</source>
            <volume>35</volume>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B4">
        <label>4.</label>
        <citation-alternatives>
          <mixed-citation publication-type="book">Tutz, G. (2011) Regression for Categorical Data. Cambridge University Press. https://doi.org/10.1017/cbo9780511842061 <pub-id pub-id-type="doi">10.1017/cbo9780511842061</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1017/cbo9780511842061">https://doi.org/10.1017/cbo9780511842061</ext-link></mixed-citation>
          <element-citation publication-type="book">
            <person-group person-group-type="author">
              <string-name>Tutz, G.</string-name>
            </person-group>
            <year>2011</year>
            <article-title>Regression for Categorical Data</article-title>
            <pub-id pub-id-type="doi">10.1017/cbo9780511842061</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B5">
        <label>5.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Agresti, A. (2010) Analysis of Ordinal Categorical Data. Wiley. https://doi.org/10.1002/9780470594001 <pub-id pub-id-type="doi">10.1002/9780470594001</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/9780470594001">https://doi.org/10.1002/9780470594001</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Agresti, A.</string-name>
            </person-group>
            <year>2010</year>
            <article-title>Analysis of Ordinal Categorical Data</article-title>
            <pub-id pub-id-type="doi">10.1002/9780470594001</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B6">
        <label>6.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Tutz, G. and Berger, M. (2017) Separating Location and Dispersion in Ordinal Regression Models. <italic>Econometrics</italic><italic>and</italic><italic>Statistics</italic>, 2, 131-148. https://doi.org/10.1016/j.ecosta.2016.10.002 <pub-id pub-id-type="doi">10.1016/j.ecosta.2016.10.002</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ecosta.2016.10.002">https://doi.org/10.1016/j.ecosta.2016.10.002</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Tutz, G.</string-name>
              <string-name>Berger, M.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Separating Location and Dispersion in Ordinal Regression Models</article-title>
            <source>Econometrics and Statistics</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.1016/j.ecosta.2016.10.002</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B7">
        <label>7.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Arvelo, I. and Plantinga, A. (2023) U.S. Mental Health Dashboard. <italic>The</italic><italic>New</italic><italic>England</italic><italic>Journal</italic><italic>of</italic><italic>Statistics</italic><italic>in</italic><italic>Data</italic><italic>Science</italic>, 2, 323-329. https://doi.org/10.51387/23-nejsds52 <pub-id pub-id-type="doi">10.51387/23-nejsds52</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.51387/23-nejsds52">https://doi.org/10.51387/23-nejsds52</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Arvelo, I.</string-name>
              <string-name>Plantinga, A.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>U</article-title>
            <source>S. Mental Health Dashboard. The New England Journal of Statistics in Data Science</source>
            <volume>2</volume>
            <pub-id pub-id-type="doi">10.51387/23-nejsds52</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B8">
        <label>8.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Fischer, M.J. (2013) Generalized Hyperbolic Secant Distributions: With Applications to Finance. Springer Science &amp; Business Media.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Fischer, M.J.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Generalized Hyperbolic Secant Distributions: With Applications to Finance</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B9">
        <label>9.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Palmitesta, P. and Provasi, C. (2004) GARCH-Type Models with Generalized Secant Hyperbolic Innovations. <italic>Studies</italic><italic>in</italic><italic>Nonlinear</italic><italic>Dynamics</italic><italic>&amp;</italic><italic>Econometrics</italic>, 8, Article 7. https://doi.org/10.2202/1558-3708.1212 <pub-id pub-id-type="doi">10.2202/1558-3708.1212</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2202/1558-3708.1212">https://doi.org/10.2202/1558-3708.1212</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Palmitesta, P.</string-name>
              <string-name>Provasi, C.</string-name>
            </person-group>
            <year>2004</year>
            <article-title>GARCH-Type Models with Generalized Secant Hyperbolic Innovations</article-title>
            <source>Studies in Nonlinear Dynamics &amp; Econometrics</source>
            <volume>8</volume>
            <elocation-id>7</elocation-id>
            <pub-id pub-id-type="doi">10.2202/1558-3708.1212</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B10">
        <label>10.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Jing, H., Liu, Y. and Zhao, J. (2022) Asymmetric Laplace Distribution Models for Financial Data: Var and Cvar. <italic>Symmetry</italic>, 14, Article 807. https://doi.org/10.3390/sym14040807 <pub-id pub-id-type="doi">10.3390/sym14040807</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/sym14040807">https://doi.org/10.3390/sym14040807</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Jing, H.</string-name>
              <string-name>Liu, Y.</string-name>
              <string-name>Zhao, J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Asymmetric Laplace Distribution Models for Financial Data: Var and Cvar</article-title>
            <source>Symmetry</source>
            <volume>14</volume>
            <elocation-id>807</elocation-id>
            <pub-id pub-id-type="doi">10.3390/sym14040807</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B11">
        <label>11.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Nolan, J.P. (2013) Financial Modeling with Heavy-Tailed Stable Distributions. <italic>WIREs</italic><italic>Computational</italic><italic>Statistics</italic>, 6, 45-55. https://doi.org/10.1002/wics.1286 <pub-id pub-id-type="doi">10.1002/wics.1286</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1002/wics.1286">https://doi.org/10.1002/wics.1286</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Nolan, J.P.</string-name>
            </person-group>
            <year>2013</year>
            <article-title>Financial Modeling with Heavy-Tailed Stable Distributions</article-title>
            <source>WIREs Computational Statistics</source>
            <volume>6</volume>
            <pub-id pub-id-type="doi">10.1002/wics.1286</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B12">
        <label>12.</label>
        <citation-alternatives>
          <mixed-citation publication-type="web">Zablotsky, B., Weeks, J.D., Terlizzi, E.P., Madans, J.H. and Blumberg, S.J. (2022) Assessing Anxiety and Depression: A Comparison of National Health Interview Survey Measures. https://stacks.cdc.gov/view/cdc/117491</mixed-citation>
          <element-citation publication-type="web">
            <person-group person-group-type="author">
              <string-name>Zablotsky, B.</string-name>
              <string-name>Weeks, J.D.</string-name>
              <string-name>Terlizzi, E.P.</string-name>
              <string-name>Madans, J.H.</string-name>
              <string-name>Blumberg, S.J.</string-name>
            </person-group>
            <year>2022</year>
            <article-title>Assessing Anxiety and Depression: A Comparison of National Health Interview Survey Measures</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B13">
        <label>13.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Terlizzi, E. and Zablotsky, B. (2024) Symptoms of Anxiety and Depression among Adults: United States, 2019 and 2022. <italic>National Health Statistics Reports</italic>, No. 213, CS353885.</mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Terlizzi, E.</string-name>
              <string-name>Zablotsky, B.</string-name>
              <string-name>Reports, N</string-name>
            </person-group>
            <year>2024</year>
            <article-title>Symptoms of Anxiety and Depression among Adults: United States, 2019 and 2022</article-title>
            <source>National Health Statistics Reports</source>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B14">
        <label>14.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Kingma, D.P. and Ba, J. (2014) Adam: A Method for Stochastic Optimization. arXiv: 1412.6980.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Kingma, D.P.</string-name>
              <string-name>Ba, J.</string-name>
            </person-group>
            <year>2014</year>
            <article-title>Adam: A Method for Stochastic Optimization</article-title>
            <fpage>1412</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B15">
        <label>15.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Williams, A.A.A. (2019) Ordinal Outcome Modeling: The Application of the Adaptive Moment Estimation Optimizer to the Elastic Net Penalized Stereotype Logit. <italic>Journal</italic><italic>of</italic><italic>Data</italic><italic>Analysis</italic><italic>and</italic><italic>Information</italic><italic>Processing</italic>, 7, 14-27. https://doi.org/10.4236/jdaip.2019.71002 <pub-id pub-id-type="doi">10.4236/jdaip.2019.71002</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4236/jdaip.2019.71002">https://doi.org/10.4236/jdaip.2019.71002</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Williams, A.A.A.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Ordinal Outcome Modeling: The Application of the Adaptive Moment Estimation Optimizer to the Elastic Net Penalized Stereotype Logit</article-title>
            <source>Journal of Data Analysis and Information Processing</source>
            <volume>7</volume>
            <pub-id pub-id-type="doi">10.4236/jdaip.2019.71002</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B16">
        <label>16.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Zablotsky, B., Lessem, S.E., Gindi, R.M., Maitland, A.K., Dahlhamer, J.M. and Blumberg, S.J. (2023) Overview of the 2019 National Health Interview Survey Questionnaire Redesign. <italic>American</italic><italic>Journal</italic><italic>of</italic><italic>Public</italic><italic>Health</italic>, 113, 408-415. https://doi.org/10.2105/ajph.2022.307197 <pub-id pub-id-type="doi">10.2105/ajph.2022.307197</pub-id><pub-id pub-id-type="pmid">36758202</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.2105/ajph.2022.307197">https://doi.org/10.2105/ajph.2022.307197</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Zablotsky, B.</string-name>
              <string-name>Lessem, S.E.</string-name>
              <string-name>Gindi, R.M.</string-name>
              <string-name>Maitland, A.K.</string-name>
              <string-name>Dahlhamer, J.M.</string-name>
              <string-name>Blumberg, S.J.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Overview of the 2019 National Health Interview Survey Questionnaire Redesign</article-title>
            <source>American Journal of Public Health</source>
            <volume>113</volume>
            <pub-id pub-id-type="doi">10.2105/ajph.2022.307197</pub-id>
            <pub-id pub-id-type="pmid">36758202</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B17">
        <label>17.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Arias de la Torre, J., Vilagut, G., Ronaldson, A., Dregan, A., Ricci-Cabello, I., Hatch, S.L., <italic>et al.</italic> (2021) Prevalence and Age Patterns of Depression in the United Kingdom. A Population-Based Study. <italic>Journal</italic><italic>of</italic><italic>Affective</italic><italic>Disorders</italic>, 279, 164-172. https://doi.org/10.1016/j.jad.2020.09.129 <pub-id pub-id-type="doi">10.1016/j.jad.2020.09.129</pub-id><pub-id pub-id-type="pmid">33059219</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.jad.2020.09.129">https://doi.org/10.1016/j.jad.2020.09.129</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Torre, J.</string-name>
              <string-name>Vilagut, G.</string-name>
              <string-name>Ronaldson, A.</string-name>
              <string-name>Dregan, A.</string-name>
              <string-name>Ricci-Cabello, I.</string-name>
              <string-name>Hatch, S.L.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Prevalence and Age Patterns of Depression in the United Kingdom</article-title>
            <source>A Population-Based Study. Journal of Affective Disorders</source>
            <volume>279</volume>
            <pub-id pub-id-type="doi">10.1016/j.jad.2020.09.129</pub-id>
            <pub-id pub-id-type="pmid">33059219</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B18">
        <label>18.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ma, F. (2021) Diagnostic and Statistical Manual of Mental Disorders-5 (DSM-5). In: u, D. and Dupre, M.E., Eds., <italic>Encyclopedia of Gerontology and Population Aging</italic>, Springer International Publishing, 1414-1425. https://doi.org/10.1007/978-3-030-22009-9_419 <pub-id pub-id-type="doi">10.1007/978-3-030-22009-9_419</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/978-3-030-22009-9_419">https://doi.org/10.1007/978-3-030-22009-9_419</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ma, F.</string-name>
              <string-name>Dupre, M.E.</string-name>
              <string-name>Aging, S</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Diagnostic and Statistical Manual of Mental Disorders-5 (DSM-5)</article-title>
            <source>In: u</source>
            <volume>1414</volume>
            <pub-id pub-id-type="doi">10.1007/978-3-030-22009-9_419</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B19">
        <label>19.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Ajele, K.W. and Idemudia, E.S. (2025) Charting the Course of Depression Care: A Meta-Analysis of Reliability Generalization of the Patient Health Questionnaire (PHQ-9) as the Measure. <italic>Discover</italic><italic>Mental</italic><italic>Health</italic>, 5, Article No. 50. https://doi.org/10.1007/s44192-025-00181-x <pub-id pub-id-type="doi">10.1007/s44192-025-00181-x</pub-id><pub-id pub-id-type="pmid">40195248</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s44192-025-00181-x">https://doi.org/10.1007/s44192-025-00181-x</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Ajele, K.W.</string-name>
              <string-name>Idemudia, E.S.</string-name>
            </person-group>
            <year>2025</year>
            <article-title>Charting the Course of Depression Care: A Meta-Analysis of Reliability Generalization of the Patient Health Questionnaire (PHQ-9) as the Measure</article-title>
            <source>Discover Mental Health</source>
            <volume>5</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1007/s44192-025-00181-x</pub-id>
            <pub-id pub-id-type="pmid">40195248</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B20">
        <label>20.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Califf, R.M., Wong, C., Doraiswamy, P.M., Hong, D.S., Miller, D.P. and Mega, J.L. (2021) Importance of Social Determinants in Screening for Depression. <italic>Journal</italic><italic>of</italic><italic>General</italic><italic>Internal</italic><italic>Medicine</italic>, 37, 2736-2743. https://doi.org/10.1007/s11606-021-06957-5 <pub-id pub-id-type="doi">10.1007/s11606-021-06957-5</pub-id><pub-id pub-id-type="pmid">34405346</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1007/s11606-021-06957-5">https://doi.org/10.1007/s11606-021-06957-5</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Califf, R.M.</string-name>
              <string-name>Wong, C.</string-name>
              <string-name>Doraiswamy, P.M.</string-name>
              <string-name>Hong, D.S.</string-name>
              <string-name>Miller, D.P.</string-name>
              <string-name>Mega, J.L.</string-name>
            </person-group>
            <year>2021</year>
            <article-title>Importance of Social Determinants in Screening for Depression</article-title>
            <source>Journal of General Internal Medicine</source>
            <volume>37</volume>
            <pub-id pub-id-type="doi">10.1007/s11606-021-06957-5</pub-id>
            <pub-id pub-id-type="pmid">34405346</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B21">
        <label>21.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Opitz, J. and Burst, S. (2019) Macro F1 and Macro F1. arXiv: 1911.03347.</mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Opitz, J.</string-name>
              <string-name>Burst, S.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>Macro F1 and Macro F1</article-title>
            <fpage>1911</fpage>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B22">
        <label>22.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Wang, H. (2023) Research on the Application of Random Forest-Based Feature Selection Algorithm in Data Mining Experiments. <italic>International</italic><italic>Journal</italic><italic>of</italic><italic>Advanced</italic><italic>Computer</italic><italic>Science</italic><italic>and</italic><italic>Applications</italic>, 14, 505-518. https://doi.org/10.14569/ijacsa.2023.0141054 <pub-id pub-id-type="doi">10.14569/ijacsa.2023.0141054</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.14569/ijacsa.2023.0141054">https://doi.org/10.14569/ijacsa.2023.0141054</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Wang, H.</string-name>
            </person-group>
            <year>2023</year>
            <article-title>Research on the Application of Random Forest-Based Feature Selection Algorithm in Data Mining Experiments</article-title>
            <source>International Journal of Advanced Computer Science and Applications</source>
            <volume>14</volume>
            <pub-id pub-id-type="doi">10.14569/ijacsa.2023.0141054</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B23">
        <label>23.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">R Core Team (2025) R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing.</mixed-citation>
          <element-citation publication-type="other">
            <year>2025</year>
            <article-title>R: A Language and Environment for Statistical Computing</article-title>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B24">
        <label>24.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Singh, V., Dwivedi, S.N. and Deo, S.V.S. (2020) Ordinal Logistic Regression Model Describing Factors Associated with Extent of Nodal Involvement in Oral Cancer Patients and Its Prospective Validation. <italic>BMC</italic><italic>Medical</italic><italic>Research</italic><italic>Methodology</italic>, 20, Article No. 95. https://doi.org/10.1186/s12874-020-00985-1 <pub-id pub-id-type="doi">10.1186/s12874-020-00985-1</pub-id><pub-id pub-id-type="pmid">32336269</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1186/s12874-020-00985-1">https://doi.org/10.1186/s12874-020-00985-1</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Singh, V.</string-name>
              <string-name>Dwivedi, S.N.</string-name>
              <string-name>Deo, S.V.S.</string-name>
            </person-group>
            <year>2020</year>
            <article-title>Ordinal Logistic Regression Model Describing Factors Associated with Extent of Nodal Involvement in Oral Cancer Patients and Its Prospective Validation</article-title>
            <source>BMC Medical Research Methodology</source>
            <volume>20</volume>
            <elocation-id>No</elocation-id>
            <pub-id pub-id-type="doi">10.1186/s12874-020-00985-1</pub-id>
            <pub-id pub-id-type="pmid">32336269</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B25">
        <label>25.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Elreedy, D. and Atiya, A.F. (2019) A Comprehensive Analysis of Synthetic Minority Oversampling Technique (SMOTE) for Handling Class Imbalance. <italic>Information</italic><italic>Sciences</italic>, 505, 32-64. https://doi.org/10.1016/j.ins.2019.07.070 <pub-id pub-id-type="doi">10.1016/j.ins.2019.07.070</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ins.2019.07.070">https://doi.org/10.1016/j.ins.2019.07.070</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Elreedy, D.</string-name>
              <string-name>Atiya, A.F.</string-name>
            </person-group>
            <year>2019</year>
            <article-title>A Comprehensive Analysis of Synthetic Minority Oversampling Technique (SMOTE) for Handling Class Imbalance</article-title>
            <source>Information Sciences</source>
            <volume>505</volume>
            <pub-id pub-id-type="doi">10.1016/j.ins.2019.07.070</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B26">
        <label>26.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Alghamdi, M., Al-Mallah, M., Keteyian, S., Brawner, C., Ehrman, J. and Sakr, S. (2017) Predicting Diabetes Mellitus Using SMOTE and Ensemble Machine Learning Approach: The Henry Ford Exercise Testing (FIT) Project. <italic>PLOS</italic><italic>ONE</italic>, 12, e0179805. https://doi.org/10.1371/journal.pone.0179805 <pub-id pub-id-type="doi">10.1371/journal.pone.0179805</pub-id><pub-id pub-id-type="pmid">28738059</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1371/journal.pone.0179805">https://doi.org/10.1371/journal.pone.0179805</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Alghamdi, M.</string-name>
              <string-name>Al-Mallah, M.</string-name>
              <string-name>Keteyian, S.</string-name>
              <string-name>Brawner, C.</string-name>
              <string-name>Ehrman, J.</string-name>
              <string-name>Sakr, S.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Predicting Diabetes Mellitus Using SMOTE and Ensemble Machine Learning Approach: The Henry Ford Exercise Testing (FIT) Project</article-title>
            <source>PLOS ONE</source>
            <volume>12</volume>
            <pub-id pub-id-type="doi">10.1371/journal.pone.0179805</pub-id>
            <pub-id pub-id-type="pmid">28738059</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B27">
        <label>27.</label>
        <citation-alternatives>
          <mixed-citation publication-type="other">Zajacova, A., Huzurbazar, S. and Todd, M. (2017) Gender and the Structure of Self-Rated Health across the Adult Life Span. <italic>Social</italic><italic>Science</italic><italic>&amp;</italic><italic>Medicine</italic>, 187, 58-66. https://doi.org/10.1016/j.socscimed.2017.06.019 <pub-id pub-id-type="doi">10.1016/j.socscimed.2017.06.019</pub-id><pub-id pub-id-type="pmid">28654822</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.socscimed.2017.06.019">https://doi.org/10.1016/j.socscimed.2017.06.019</ext-link></mixed-citation>
          <element-citation publication-type="other">
            <person-group person-group-type="author">
              <string-name>Zajacova, A.</string-name>
              <string-name>Huzurbazar, S.</string-name>
              <string-name>Todd, M.</string-name>
            </person-group>
            <year>2017</year>
            <article-title>Gender and the Structure of Self-Rated Health across the Adult Life Span</article-title>
            <source>Social Science &amp; Medicine</source>
            <volume>187</volume>
            <pub-id pub-id-type="doi">10.1016/j.socscimed.2017.06.019</pub-id>
            <pub-id pub-id-type="pmid">28654822</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
      <ref id="B28">
        <label>28.</label>
        <citation-alternatives>
          <mixed-citation publication-type="journal">Bandura, A. (1997) Editorial. <italic>American</italic><italic>Journal</italic><italic>of</italic><italic>Health</italic><italic>Promotion</italic>, 12, 8-10. https://doi.org/10.4278/0890-1171-12.1.8 <pub-id pub-id-type="doi">10.4278/0890-1171-12.1.8</pub-id><pub-id pub-id-type="pmid">10170438</pub-id><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4278/0890-1171-12.1.8">https://doi.org/10.4278/0890-1171-12.1.8</ext-link></mixed-citation>
          <element-citation publication-type="journal">
            <person-group person-group-type="author">
              <string-name>Bandura, A.</string-name>
            </person-group>
            <year>1997</year>
            <article-title>Editorial</article-title>
            <source>American Journal of Health Promotion</source>
            <volume>12</volume>
            <pub-id pub-id-type="doi">10.4278/0890-1171-12.1.8</pub-id>
            <pub-id pub-id-type="pmid">10170438</pub-id>
          </element-citation>
        </citation-alternatives>
      </ref>
    </ref-list>
  </back>
</article>