<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    ojs
   </journal-id>
   <journal-title-group>
    <journal-title>
     Open Journal of Statistics
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2161-718X
   </issn>
   <issn publication-format="print">
    2161-7198
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/ojs.2025.152011
   </article-id>
   <article-id pub-id-type="publisher-id">
    ojs-142110
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Physics 
     </subject>
     <subject>
       Mathematics
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    Dynamic Conditional Feature Screening: A High-Dimensional Feature Selection Method Based on Mutual Information and Regression Error
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Yi
      </surname>
      <given-names>
       Zhao
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Guangming
      </surname>
      <given-names>
       Deng
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aSchool of Mathematics and Statistics, Guilin University of Technology, Guilin, China
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aApplied Statistics Institute, Guilin University of Technology, Guilin, China
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     31
    </day> 
    <month>
     03
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    15
   </volume> 
   <issue>
    02
   </issue>
   <fpage>
    199
   </fpage>
   <lpage>
    242
   </lpage>
   <history>
    <date date-type="received">
     <day>
      17,
     </day>
     <month>
      March
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      19,
     </day>
     <month>
      March
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      19,
     </day>
     <month>
      April
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    Current high-dimensional feature screening methods still face significant challenges in handling mixed linear and nonlinear relationships, controlling redundant information, and improving model robustness. In this study, we propose a Dynamic Conditional Feature Screening (DCFS) method tailored for high-dimensional economic forecasting tasks. Our goal is to accurately identify key variables, enhance predictive performance, and provide both theoretical foundations and practical tools for macroeconomic modeling. The DCFS method constructs a comprehensive test statistic by integrating conditional mutual information with conditional regression error differences. By introducing a dynamic weighting mechanism, DCFS adaptively balances the linear and nonlinear contributions of features during the screening process. In addition, a dynamic thresholding mechanism is designed to effectively control the false discovery rate (FDR), thereby improving the stability and reliability of the screening results. On the theoretical front, we rigorously prove that the proposed method satisfies the sure screening property and rank consistency, ensuring accurate identification of the truly important feature set in high-dimensional settings. Simulation results demonstrate that under purely linear, purely nonlinear, and mixed dependency structures, DCFS consistently outperforms classical screening methods such as SIS, CSIS, and IG-SIS in terms of true positive rate (TPR), false discovery rate (FDR), and rank correlation. These results highlight the superior accuracy, robustness, and stability of our method. Furthermore, an empirical analysis based on the U.S. FRED-MD macroeconomic dataset confirms the practical value of DCFS in real-world forecasting tasks. The experimental results show that DCFS achieves lower prediction errors (RMSE and MAE) and higher R
    <sup>2</sup> values in forecasting GDP growth. The selected key variables—including the Industrial Production Index (IP), Federal Funds Rate, Consumer Price Index (CPI), and Money Supply (M2)—possess clear economic interpretability, offering reliable support for economic forecasting and policy formulation.
   </abstract>
   <kwd-group> 
    <kwd>
     High-Dimensional Feature Screening
    </kwd> 
    <kwd>
      Conditional Mutual Information
    </kwd> 
    <kwd>
      Regression Error Difference
    </kwd> 
    <kwd>
      Dynamic Weighting
    </kwd> 
    <kwd>
      Dynamic Thresholding
    </kwd> 
    <kwd>
      Macroeconomic Forecasting
    </kwd> 
    <kwd>
      FRED-MD Dataset
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>With the advancement of data collection technology, high-dimensional data has been widely applied in various fields such as bioinformatics, financial analysis, environmental science, and medical diagnosis. A significant feature of high-dimensional data is that the dimensionality p of covariates is greater than the sample size n. Specifically, when the dimensionality p of the data is much larger than the sample size n, this type of data is often referred to as ultra-high-dimensional data. When analyzing ultra-high dimensional data, it is often assumed that there are fewer covariates that affect the response variable, which is referred to as the "sparsity" hypothesis in the literature. Under the assumption of sparsity, determining the covariates that truly affect the response variable has become an important and fundamental issue. Fan and L ü (2008) <xref ref-type="bibr" rid="scirp.142110-1">
     [1]
    </xref> proposed a method for screening variables based on Pearson correlation coefficient and named it Sure Independence Screening (SIS). Inspired by Fan and L ü (2008), the study of variable selection methods has received great attention from statisticians, resulting in a large number of research achievements. Li, Zhong, and Zhu (2012) <xref ref-type="bibr" rid="scirp.142110-2">
     [2]
    </xref> proposed a deterministic independence screening method using distance correlation. According to Li et al. (2012), a screening method based on distance correlation coefficient was proposed. Shao and Zhang (2014) <xref ref-type="bibr" rid="scirp.142110-3">
     [3]
    </xref> proposed the martingale difference correlation coefficient screening method (MDC), which measures the deviation between two random variables and one correlation coefficient. Mai and Zou (2015) <xref ref-type="bibr" rid="scirp.142110-4">
     [4]
    </xref> developed a binary classification model screening method based on Kolmogorov distance. Ni and Fang (2016) <xref ref-type="bibr" rid="scirp.142110-5">
     [5]
    </xref> constructed corresponding statistical measures for correlation analysis from the perspective of information content, and proposed a method for selecting ultra-high dimensional variables based on information gain (IG-SIS). Zhu Yidan and Chen Xingrong (2021) <xref ref-type="bibr" rid="scirp.142110-6">
     [6]
    </xref> proposed a variable selection method based on information gain rate (IGR-SIS), which has higher accuracy in screening important variables compared to IG-SIS. Fan et al. (2020) <xref ref-type="bibr" rid="scirp.142110-7">
     [7]
    </xref> and Zeng Jin and Zhou Jianjun (2017) <xref ref-type="bibr" rid="scirp.142110-8">
     [8]
    </xref> provided a comprehensive overview of variable selection and feature screening.</p>
   <p>Feature screening methods based on marginal correlations, such as MDC <xref ref-type="bibr" rid="scirp.142110-3">
     [3]
    </xref> and IG-SIS <xref ref-type="bibr" rid="scirp.142110-5">
     [5]
    </xref>, often fail to capture complex conditional dependencies among variables. To address this challenge, Fan et al. (2016) <xref ref-type="bibr" rid="scirp.142110-9">
     [9]
    </xref> proposed the Conditional Sure Independence Screening (CSIS) method, which identifies important features by estimating conditional distributions or regression functions. This method reduces the impact of redundant variables while ensuring feature saliency. Research has shown that CSIS outperforms traditional SIS methods in processing data with nonlinear dependency structures. Lin et al. (2020) <xref ref-type="bibr" rid="scirp.142110-10">
     [10]
    </xref> Propose a model free conditional feature selection method based on conditional distance correlation. Zhou et al. <xref ref-type="bibr" rid="scirp.142110-11">
     [11]
    </xref> proposed conditional feature selection for variable coefficient models. Xiong et al. <xref ref-type="bibr" rid="scirp.142110-12">
     [12]
    </xref> proposed a new model free interaction screening program called MCVIS by introducing the MCV index to quantify the importance of the interaction effects between predictive factors. Wang et al. (2023) <xref ref-type="bibr" rid="scirp.142110-13">
     [13]
    </xref> proposed a new model free conditional screening method using conditional feature functions as screening criteria for massive imbalanced data.</p>
   <p>Although methods such as SIS <xref ref-type="bibr" rid="scirp.142110-1">
     [1]
    </xref>, CSIS <xref ref-type="bibr" rid="scirp.142110-9">
     [9]
    </xref>, and IG-SIS <xref ref-type="bibr" rid="scirp.142110-5">
     [5]
    </xref> are effective in certain scenarios, they often overlook complex conditional dependencies among variables. As a result, they struggle to simultaneously capture both linear and nonlinear features and tend to exhibit poor screening stability when dealing with high-dimensional data. Some existing methods have attempted to incorporate conditional screening mechanisms, yet still suffer from insufficient characterization of nonlinear structures and limited control over screening errors. To address these challenges, this paper proposes a Dynamic Conditional Feature Screening (DCFS) method tailored to high-dimensional economic forecasting tasks. The specific goals are: (1) to design a hybrid test statistic that combines conditional mutual information and regression error difference for a more comprehensive assessment of feature importance; (2) to develop dynamic weighting and thresholding mechanisms to enhance the model’s adaptability to complex data structures and improve screening robustness; (3) to establish the theoretical guarantees of the proposed method, including sure screening and consistency; and (4) to validate its effectiveness through both simulation studies and empirical analysis using real macroeconomic data. This work aims to provide a novel methodology that offers both theoretical rigor and practical value for high-dimensional data analysis, feature selection, and macroeconomic modeling.</p>
   <p>The structure of this paper is as follows: Section 2 introduces the mathematical framework of DCFS; Section 3 presents its theoretical properties; Section 4 conducts simulation experiments to evaluate empirical performance; Section 5 applies DCFS to real-world macroeconomic forecasting tasks; and Section 6 concludes the paper and discusses future research directions.</p>
  </sec><sec id="s2">
   <title>2. Dynamic Conditional Feature Selection Method</title>
   <sec id="s2_1">
    <title>2.1. Variable Setting and Objectives</title>
    <p>Variable setting: Let the predictor variables be denoted by</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         X 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            p 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ∈ 
       </mo> 
       <msup> 
        <mi>
          ℝ 
        </mi> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           × 
         </mo> 
         <mi>
           p 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (1)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math> is the dimensionality of the predictors, and each 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> represents a candidate feature variable. The conditional variables are defined as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Z 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            Z 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            Z 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            Z 
          </mi> 
          <mi>
            q 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ∈ 
       </mo> 
       <msup> 
        <mi>
          ℝ 
        </mi> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           × 
         </mo> 
         <mi>
           q 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (2)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        q 
      </mi> 
     </math> is the number of conditional variables, which are known to be strongly associated with the response variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         Y 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <msup> 
        <mi>
          ℝ 
        </mi> 
        <mi>
          n 
        </mi> 
       </msup> 
      </mrow> 
     </math>, serving as the target variable. Such associations between 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> may arise from the inherent structure of the data or prior domain knowledge. The conditional variables play a critical role in adjusting for and eliminating confounding effects during the screening process, thus enabling a more accurate identification of the true relationship between the predictors and the response variable. The full dataset consists of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        n 
      </mi> 
     </math> independent observations denoted by 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mrow> 
         <mrow> 
          <mo>
            { 
          </mo> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                X 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
             <mo>
               , 
             </mo> 
             <msub> 
              <mi>
                Z 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
             <mo>
               , 
             </mo> 
             <msub> 
              <mi>
                Y 
              </mi> 
              <mi>
                i 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            } 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          n 
        </mi> 
       </msubsup> 
      </mrow> 
     </math>.</p>
    <p>Screening objective: Given the conditional variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>, our objective is to identify a subset 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         S 
       </mi> 
       <mo>
         ⊆ 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           2 
         </mn> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <mi>
           p 
         </mi> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math> of predictors that make a significant contribution to the response variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>. To this end, we define the set of active predictors as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         D 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           : 
         </mo> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <mo>
           ⊥ 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math> and the set of inactive predictors as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mi>
           k 
         </mi> 
         <mo>
           : 
         </mo> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            k 
          </mi> 
         </msub> 
         <mo>
           ⊥ 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. Our goal is to accurately recover the active predictor set 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi mathvariant="script">
        D 
      </mi> 
     </math>, that is, to identify all features 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> satisfying 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ⊥ 
       </mo> 
       <mi>
         Y 
       </mi> 
       <mo>
         | 
       </mo> 
       <mtext>
         Z 
       </mtext> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mo>
         ∀ 
       </mo> 
       <mi>
         j 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mi>
         S 
       </mi> 
      </mrow> 
     </math>. The screening procedure should guarantee high identification accuracy of truly important predictors while minimizing the risk of omission.</p>
   </sec>
   <sec id="s2_2">
    <title>2.2. Construction of Conditional Correlation Measurement</title>
    <p>Conditional mutual information 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is employed to quantify the dependency between the feature variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> and the response variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>, given the conditional variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>. It serves as a core component in the construction of the dynamic test statistic proposed in this study and is particularly well-suited for capturing nonlinear relationships.</p>
    <p>The definition of conditional mutual information is based on the difference in conditional entropies, expressed as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         H 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         − 
       </mo> 
       <mi>
         H 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (3)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         H 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mo>
          ⋅ 
        </mo> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> denotes the conditional entropy. The two terms in Equation (3) can be further expressed as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         H 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mo>
         − 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           log 
         </mi> 
         <msub> 
          <mi>
            f 
          </mi> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              x 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             | 
           </mo> 
           <mtext>
             z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mi>
         H 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mo>
         − 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           log 
         </mi> 
         <msub> 
          <mi>
            f 
          </mi> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             | 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             , 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              x 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             | 
           </mo> 
           <mi>
             y 
           </mi> 
           <mo>
             , 
           </mo> 
           <mtext>
             z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (4)</p>
    <p>By substituting Equation (4) into Equation (3), the conditional mutual information can be written as a log-likelihood ratio:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           log 
         </mi> 
         <mfrac> 
          <mrow> 
           <msub> 
            <mi>
              f 
            </mi> 
            <mrow> 
             <msub> 
              <mi>
                X 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
             <mo>
               | 
             </mo> 
             <mtext>
               Z 
             </mtext> 
            </mrow> 
           </msub> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                x 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
             <mo>
               | 
             </mo> 
             <mtext>
               z 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mrow> 
           <msub> 
            <mi>
              f 
            </mi> 
            <mrow> 
             <msub> 
              <mi>
                X 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
             <mo>
               | 
             </mo> 
             <mi>
               Y 
             </mi> 
             <mo>
               , 
             </mo> 
             <mtext>
               Z 
             </mtext> 
            </mrow> 
           </msub> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                x 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
             <mo>
               | 
             </mo> 
             <mi>
               y 
             </mi> 
             <mo>
               , 
             </mo> 
             <mtext>
               z 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </mfrac> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (5)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> denote the conditional probability density functions.From an information-theoretic perspective, conditional mutual information measures the additional information about 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> that 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> provides when 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math> is already known. When 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> are conditionally independent given 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>, we have 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>; otherwise, if 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> provides independent information about 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>, then 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>.</p>
    <p>Compared to traditional linear correlation coefficients or least squares estimates, conditional mutual information offers several advantages:</p>
    <p>It captures nonlinear, non-symmetric, and complex interaction relationships;</p>
    <p>In the proposed Dynamic Conditional Feature Screening (DCFS) method, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is combined with the conditional regression error difference 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> to construct a weighted test statistic (see Section 2.3). These two components jointly measure variable contributions from nonlinear and linear perspectives, respectively, providing a more comprehensive and adaptive basis for high-dimensional feature screening.</p>
    <p>The conditional regression error difference, denoted as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, measures the incremental explanatory power of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> for the response variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>. It is defined as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         Δ 
       </mi> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <msup> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               Y 
             </mi> 
             <mo>
               − 
             </mo> 
             <msup> 
              <mover accent="true"> 
               <mi>
                 m 
               </mi> 
               <mo>
                 ^ 
               </mo> 
              </mover> 
              <mi>
                j 
              </mi> 
             </msup> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mtext>
                Z 
              </mtext> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         − 
       </mo> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <msup> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               Y 
             </mi> 
             <mo>
               − 
             </mo> 
             <msup> 
              <mover accent="true"> 
               <mi>
                 m 
               </mi> 
               <mo>
                 ^ 
               </mo> 
              </mover> 
              <mrow> 
               <mrow> 
                <mo>
                  ( 
                </mo> 
                <mrow> 
                 <msub> 
                  <mi>
                    X 
                  </mi> 
                  <mi>
                    j 
                  </mi> 
                 </msub> 
                 <mo>
                   , 
                 </mo> 
                 <mtext>
                   Z 
                 </mtext> 
                </mrow> 
                <mo>
                  ) 
                </mo> 
               </mrow> 
              </mrow> 
             </msup> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  X 
                </mi> 
                <mi>
                  j 
                </mi> 
               </msub> 
               <mo>
                 , 
               </mo> 
               <mtext>
                 Z 
               </mtext> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (6)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mover accent="true"> 
         <mi>
           m 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mtext>
          Z 
        </mtext> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is the regression model based solely on 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mover accent="true"> 
         <mi>
           m 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is the regression model based on both 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>. A larger 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
      </mrow> 
     </math> indicates stronger incremental explanatory ability of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> with respect to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>. This quantity focuses on assessing the contribution of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> from a regression analysis perspective, effectively reflecting its utility in explaining variations in the response variable.</p>
   </sec>
   <sec id="s2_3">
    <title>2.3. Dynamic Weight Control Mechanism</title>
    <p>To account for potential linear and nonlinear dependencies between features and the response, the proposed Dynamic Conditional Feature Screening (DCFS) method introduces a novel dynamic weighting mechanism to enhance adaptivity and robustness in the feature screening process.</p>
    <p>The core idea of this mechanism is to dynamically adjust the relative weights of linear and nonlinear contributions for each feature 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math>, based on its statistical dependence with the response 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>, conditional on 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>. Two complementary sources of information are considered in DCFS method: conditional mutual information 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> and conditional regression error difference 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> measure the nonlinear dependence and linear predictive power between the characteristic and response variables, respectively.</p>
    <p>Specifically, we construct the following weighted statistic to evaluate the importance of each feature:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mi>
         Δ 
       </mi> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (7)</p>
    <p>where the weights 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> are derived from the relative information contribution of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math>, defined as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mi>
           I 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             ; 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           I 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             ; 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mtext>
           Δ 
         </mtext> 
         <mi>
           E 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </mfrac> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Δ 
         </mtext> 
         <mi>
           E 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           I 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             ; 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mtext>
           Δ 
         </mtext> 
         <mi>
           E 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mi>
             Y 
           </mi> 
           <mo>
             | 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </mfrac> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (8)</p>
    <p>This weighting mechanism enables automatic adjustment based on the specific information structure of each feature. When 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≫ 
       </mo> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, indicating dominant nonlinear information, the weight tends toward 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         → 
       </mo> 
       <mn>
         1 
       </mn> 
      </mrow> 
     </math>, and the statistic emphasizes nonlinear dependence. Conversely, when 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≫ 
       </mo> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, the statistic primarily reflects linear predictive ability.</p>
    <p>In terms of theoretical properties, the statistic 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> has good asymptotic properties under large sample conditions:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mi>
          p 
        </mi> 
       </mover> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mn>
          0 
        </mn> 
       </msubsup> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         as 
       </mtext> 
       <mtext>
           
       </mtext> 
       <mi>
         n 
       </mi> 
       <mo>
         → 
       </mo> 
       <mi>
         ∞ 
       </mi> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (9)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mn>
          0 
        </mn> 
       </msubsup> 
      </mrow> 
     </math> denotes the true information content, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mi>
          p 
        </mi> 
       </mover> 
      </mrow> 
     </math> denotes convergence in probability.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msqrt> 
        <mi>
          n 
        </mi> 
       </msqrt> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
          <mrow> 
           <mtext>
             dynamic 
           </mtext> 
          </mrow> 
         </msubsup> 
         <mo>
           − 
         </mo> 
         <msubsup> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
          <mn>
            0 
          </mn> 
         </msubsup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mi>
          d 
        </mi> 
       </mover> 
       <mi>
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            σ 
          </mi> 
          <mi>
            j 
          </mi> 
          <mn>
            2 
          </mn> 
         </msubsup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (10)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mi>
          d 
        </mi> 
       </mover> 
      </mrow> 
     </math> denotes convergence in distribution, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          σ 
        </mi> 
        <mi>
          j 
        </mi> 
        <mn>
          2 
        </mn> 
       </msubsup> 
      </mrow> 
     </math> is the asymptotic variance.</p>
    <p>These asymptotic properties provide a solid theoretical foundation for establishing the sure screening property and rank consistency, which are discussed in subsequent sections.</p>
    <p>In summary, by incorporating a dynamic weighting mechanism, the DCFS method allows the assessment of feature importance to adaptively respond to both linear and nonlinear structures in the data. This enables more accurate and robust evaluation of variable contributions under various dependency scenarios, significantly enhancing the adaptability and predictive performance of high-dimensional feature screening.</p>
   </sec>
   <sec id="s2_4">
    <title>2.4. Statistical Estimation Methods</title>
    <p>According to formula (5) in section 2.2.1, the Accurately estimating the conditional mutual information requires reliable estimation of the conditional densities 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>. Traditional approaches such as kernel methods face several limitations in high-dimensional settings. Specifically, they require careful selection of kernel functions and bandwidth parameters, which often involve extensive manual tuning and domain knowledge. As data dimensionality increases, the computational complexity of kernel methods grows exponentially, making them inefficient for large-scale problems. Similarly, K-nearest neighbor (KNN) methods involve pairwise distance computations among all samples, which results in significant computational overhead and memory consumption when applied to large datasets.</p>
    <p>To overcome these challenges, we employ the Variational Autoencoder (VAE) <xref ref-type="bibr" rid="scirp.142110-14">
      [14]
     </xref> framework from deep learning to estimate the conditional distributions. The VAE approximates</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≈ 
       </mo> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          θ 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≈ 
       </mo> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          θ 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (11)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        θ 
      </mi> 
     </math> denotes the parameters of a neural network that models the conditional distributions. The VAE architecture consists of two main components: an encoder and a decoder. The encoder takes 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math> as input and outputs the latent mean and variance, defining the approximate posterior distribution as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          q 
        </mi> 
        <mi>
          ϕ 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           h 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="script">
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            μ 
          </mi> 
          <mi>
            ϕ 
          </mi> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mtext>
            z 
          </mtext> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mtext>
           diag 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msubsup> 
            <mi>
              σ 
            </mi> 
            <mi>
              ϕ 
            </mi> 
            <mn>
              2 
            </mn> 
           </msubsup> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mtext>
              z 
            </mtext> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (12)</p>
    <p>which maps the input condition 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math> into a latent space, extracting essential features. The decoder then reconstructs the target variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> from the latent representation 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        h 
      </mi> 
     </math>, modeling the conditional distribution as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          θ 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="script">
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            μ 
          </mi> 
          <mi>
            θ 
          </mi> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             h 
           </mi> 
           <mo>
             , 
           </mo> 
           <mtext>
             z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            σ 
          </mi> 
          <mi>
            θ 
          </mi> 
          <mn>
            2 
          </mn> 
         </msubsup> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             h 
           </mi> 
           <mo>
             , 
           </mo> 
           <mtext>
             z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (13)</p>
    <p>thereby learning the mapping between the latent variables and the observed target.</p>
    <p>In our implementation, the VAE network used for estimating conditional mutual information has the following architecture: the input layer takes features of dimension 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math>; the encoder consists of two hidden layers, each with 64 neurons and ReLU activation functions; the latent space has a dimension of 32; and the decoder mirrors the encoder structure with two hidden layers and ReLU activations.</p>
    <p>For model training, we adopt the Adam optimizer with a learning rate of 0.001 and a batch size of 128. The maximum number of training epochs is set to 500. To prevent overfitting and reduce unnecessary computation, we apply an early stopping strategy that halts training if the loss does not improve for 50 consecutive epochs. Additionally, L2 regularization with a penalty coefficient 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         λ 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         0.0001 
       </mn> 
      </mrow> 
     </math> is employed to improve generalization performance.</p>
    <p>To estimate the conditional regression error difference, we use two separate Multilayer Perceptron (MLP) <xref ref-type="bibr" rid="scirp.142110-15">
      [15]
     </xref> models to approximate the regression functions 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          m 
        </mi> 
        <mi>
          j 
        </mi> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mtext>
          Z 
        </mtext> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msup> 
        <mi>
          m 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msup> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, respectively.</p>
    <p>MLP is a powerful class of neural networks composed of multiple hidden layers, capable of learning complex features and patterns from the input data. During training, the MLPs iteratively update their weights and biases using a large number of training samples, so that their outputs closely approximate the true regression functions. Once the models are trained, we compute the conditional regression error difference as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mtext>
         MSE 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mover accent="true"> 
           <mi>
             Y 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mtext>
            Z 
          </mtext> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         − 
       </mo> 
       <mtext>
         MSE 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mover accent="true"> 
           <mi>
             Y 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                X 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
             <mo>
               , 
             </mo> 
             <mtext>
               Z 
             </mtext> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (14)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           Y 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mtext>
          Z 
        </mtext> 
       </msub> 
      </mrow> 
     </math> denotes the prediction of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> based on 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           Y 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mtext>
             Z 
           </mtext> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is the prediction based on 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>.This MLP-based estimation approach effectively leverages the representational capacity of neural networks to capture both linear and nonlinear dependencies between the feature and the response variable, resulting in an accurate estimation of the conditional regression error difference.</p>
    <p>In our implementation, both MLP models share the same architecture: each contains two hidden layers with 64 neurons per layer and uses the ReLU activation function. The training configuration is consistent with that of the VAE model, employing the Adam optimizer with a learning rate of 0.001, a batch size of 128, and a maximum of 500 training epochs. An early stopping strategy is applied to terminate training if the loss does not improve for 50 consecutive epochs.</p>
    <p>To comprehensively evaluate the practical efficiency of the proposed Dynamic Conditional Feature Screening (DCFS) method, this section analyzes its computational complexity and compares its runtime performance with several classical feature screening methods, including SIS, CSIS, DC-SIS, and IG-SIS.</p>
    <p>(1) Time Complexity Analysis</p>
    <p>The computational complexity of DCFS primarily arises from two key components: conditional mutual information estimation and conditional regression error difference estimation. For conditional mutual information, we adopt a Variational Autoencoder (VAE)-based estimation approach. The computational cost per training iteration depends on the network architecture and training process. In our implementation, the VAE consists of two hidden layers with 64 neurons each. The training complexity is approximately 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi mathvariant="script">
         O 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <msup> 
          <mi>
            p 
          </mi> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        n 
      </mi> 
     </math> is the sample size and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math> is the number of features. For conditional regression error difference estimation, we use a Multilayer Perceptron (MLP), which has a similar computational complexity of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi mathvariant="script">
         O 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <msup> 
          <mi>
            p 
          </mi> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. This is because both forward propagation and backpropagation involve operations over all feature variables and require learning complex interactions.</p>
    <p>Therefore, the overall time complexity of the DCFS method can be approximated as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi mathvariant="script">
         O 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <msup> 
          <mi>
            p 
          </mi> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>.</p>
    <p>(2) Space Complexity Analysis</p>
    <p>In terms of space complexity, DCFS requires storing the original data matrix, intermediate parameters, and results during the estimation of conditional mutual information and regression error difference. Hence, the overall space complexity is approximately 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi mathvariant="script">
         O 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mi>
           p 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. As the number of features increases, the space usage grows linearly, which aligns well with typical memory constraints in high-dimensional data environments.</p>
    <p>(3) Benchmark Runtime Comparison</p>
    <p>To further quantify the practical runtime performance of DCFS, we conduct benchmark experiments under three feature dimensionality settings: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         p 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         500 
       </mn> 
      </mrow> 
     </math>, 1000, and 2000. We compare the runtime (in seconds) of DCFS with SIS, CSIS, DC-SIS, and IG-SIS. The results are summarized in <xref ref-type="table" rid="table1">
      Table 1
     </xref>:</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 1. Benchmark runtime comparison.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="28.90%"><p style="text-align:center">Feature Dimension p</p></td> 
       <td class="custom-bottom-td acenter" width="10.75%"><p style="text-align:center">SIS (s)</p></td> 
       <td class="custom-bottom-td acenter" width="15.10%"><p style="text-align:center">CSIS (s)</p></td> 
       <td class="custom-bottom-td acenter" width="16.80%"><p style="text-align:center">DC-SIS (s)</p></td> 
       <td class="custom-bottom-td acenter" width="14.22%"><p style="text-align:center">IG-SIS (s)</p></td> 
       <td class="custom-bottom-td acenter" width="14.23%"><p style="text-align:center">DCFS (s)</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="28.90%"><p style="text-align:center">500</p></td> 
       <td class="custom-top-td acenter" width="10.75%"><p style="text-align:center">0.52</p></td> 
       <td class="custom-top-td acenter" width="15.10%"><p style="text-align:center">1.35</p></td> 
       <td class="custom-top-td acenter" width="16.80%"><p style="text-align:center">3.42</p></td> 
       <td class="custom-top-td acenter" width="14.22%"><p style="text-align:center">4.68</p></td> 
       <td class="custom-top-td acenter" width="14.23%"><p style="text-align:center">5.21</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="28.90%"><p style="text-align:center">1000</p></td> 
       <td class="acenter" width="10.75%"><p style="text-align:center">1.05</p></td> 
       <td class="acenter" width="15.10%"><p style="text-align:center">2.71</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">7.58</p></td> 
       <td class="acenter" width="14.22%"><p style="text-align:center">9.84</p></td> 
       <td class="acenter" width="14.23%"><p style="text-align:center">10.32</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="28.90%"><p style="text-align:center">2000</p></td> 
       <td class="acenter" width="10.75%"><p style="text-align:center">2.31</p></td> 
       <td class="acenter" width="15.10%"><p style="text-align:center">5.87</p></td> 
       <td class="acenter" width="16.80%"><p style="text-align:center">15.67</p></td> 
       <td class="acenter" width="14.22%"><p style="text-align:center">21.56</p></td> 
       <td class="acenter" width="14.23%"><p style="text-align:center">22.14</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>As shown in the table, although DCFS requires slightly more computational time compared to classical methods, it maintains a relatively acceptable runtime performance. More importantly, in high-dimensional scenarios, DCFS exhibits a stable growth pattern in complexity that aligns well with practical application demands.</p>
    <p>In summary, the above complexity analysis and benchmark comparisons confirm that DCFS offers good scalability for large-scale data applications and is capable of supporting efficient and reliable feature screening tasks in real-world high-dimensional environments.</p>
   </sec>
   <sec id="s2_5">
    <title>2.5. False Discovery Rate Control Mechanism</title>
    <p>
     <xref ref-type="bibr" rid="scirp.142110-"></xref>To control the False Discovery Rate (FDR), we adopt the Reflection via Data Splitting (REDS) method <xref ref-type="bibr" rid="scirp.142110-16">
      [16]
     </xref> to construct a data-driven dynamic threshold 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mrow> 
         <mtext>
           threshold 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math>. The procedure is outlined as follows:</p>
    <p>Data Splitting: The original dataset is randomly divided into two disjoint subsets, denoted as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi mathvariant="script">
          D 
        </mi> 
        <mi>
          A 
        </mi> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi mathvariant="script">
          D 
        </mi> 
        <mi>
          B 
        </mi> 
       </msub> 
      </mrow> 
     </math>.</p>
    <p>Preliminary Screening: On subset 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi mathvariant="script">
          D 
        </mi> 
        <mi>
          A 
        </mi> 
       </msub> 
      </mrow> 
     </math>, we compute the dynamic importance statistic 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> for each feature. The computation strictly follows the procedures and parameter settings described in earlier sections to ensure accuracy and consistency. The resulting statistics serve as the basis for subsequent significance testing.</p>
    <p>Reflection Testing: On subset 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi mathvariant="script">
          D 
        </mi> 
        <mi>
          B 
        </mi> 
       </msub> 
      </mrow> 
     </math>, we simulate the null distribution 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           null 
         </mtext> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> under the no-signal assumption. This is achieved by shuffling the pairwise correspondence between 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        X 
      </mtext> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>, i.e.,</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           null 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         ~ 
       </mo> 
       <mtext>
         shuffle 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi mathvariant="script">
            D 
          </mi> 
          <mi>
            A 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi mathvariant="script">
            D 
          </mi> 
          <mi>
            B 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (15)</p>
    <p>which reflects a condition where no real association exists between features and the response. This provides an empirical estimate of the distribution of test statistics under the null hypothesis.</p>
    <p>Significance Thresholding: A data-adaptive threshold is determined by computing the 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           − 
         </mo> 
         <mi>
           α 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>-quantile of the null distribution:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mrow> 
         <mtext>
           threshold 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mtext>
         quantile 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
          <mrow> 
           <mtext>
             null 
           </mtext> 
          </mrow> 
         </msubsup> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           − 
         </mo> 
         <mi>
           α 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (16)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         α 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is the user-specified FDR level. This threshold ensures that the proportion of falsely selected features among all selected features satisfies</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         FDR 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Number of False Positives 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           Number of Selected Features 
         </mtext> 
        </mrow> 
       </mfrac> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         α 
       </mi> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (17)</p>
    <p>In practice, the choice of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        α 
      </mi> 
     </math> should balance domain-specific risk tolerance and the desired level of selection conservativeness.</p>
    <p>Final Selection: The final screened feature set consists of all features satisfying</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         &gt; 
       </mo> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mrow> 
         <mtext>
           threshold 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (18)</p>
    <p>These features are considered to have statistically significant influence on the response variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> and are retained for downstream analysis.</p>
   </sec>
  </sec><sec id="s3">
   <title>3. Theoretical Properties and Proof of DCFS Method</title>
   <sec id="s3_1">
    <title>3.1. Non Negativity and Distribution Irrelevance</title>
    <p>The proposed unified statistic 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math></p>
    <p>integrates the complementary strengths of conditional mutual information and conditional regression error difference, and enjoys the following theoretical guarantees:</p>
    <p>1. Non-negativity:</p>
    <p>Both components of the statistic are non-negative by definition. The conditional mutual information satisfies</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         0 
       </mn> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (19)</p>
    <p>as it measures the amount of information shared between variables and cannot be negative. Similarly, the conditional regression error difference</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <msup> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               Y 
             </mi> 
             <mo>
               − 
             </mo> 
             <msup> 
              <mi>
                m 
              </mi> 
              <mi>
                j 
              </mi> 
             </msup> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mtext>
                Z 
              </mtext> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         − 
       </mo> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <msup> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               Y 
             </mi> 
             <mo>
               − 
             </mo> 
             <msup> 
              <mi>
                m 
              </mi> 
              <mrow> 
               <mrow> 
                <mo>
                  ( 
                </mo> 
                <mrow> 
                 <msub> 
                  <mi>
                    X 
                  </mi> 
                  <mi>
                    j 
                  </mi> 
                 </msub> 
                 <mo>
                   , 
                 </mo> 
                 <mtext>
                   Z 
                 </mtext> 
                </mrow> 
                <mo>
                  ) 
                </mo> 
               </mrow> 
              </mrow> 
             </msup> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  X 
                </mi> 
                <mi>
                  j 
                </mi> 
               </msub> 
               <mo>
                 , 
               </mo> 
               <mtext>
                 Z 
               </mtext> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         0 
       </mn> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (20)</p>
    <p>because the inclusion of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> in the model does not worsen its predictive accuracy. That is, adding 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> will not increase the expected squared prediction error. Therefore, the combined statistic satisfies</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         0 
       </mn> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mo>
         ∀ 
       </mo> 
       <mi>
         j 
       </mi> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (21)</p>
    <p>ensuring that its values remain meaningful and interpretable in all cases.</p>
    <p>2. Distribution-Free Robustness:</p>
    <p>The limiting distribution of the conditional mutual information 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> under the null hypothesis 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          H 
        </mi> 
        <mn>
          0 
        </mn> 
       </msub> 
       <mo>
         : 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ⊥ 
       </mo> 
       <mi>
         Y 
       </mi> 
       <mo>
         | 
       </mo> 
       <mtext>
         Z 
       </mtext> 
      </mrow> 
     </math> is distribution-free, i.e., it does not depend on the specific joint distribution of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         , 
       </mo> 
       <mi>
         Y 
       </mi> 
       <mo>
         , 
       </mo> 
       <mtext>
         Z 
       </mtext> 
      </mrow> 
     </math>. This property enables hypothesis testing and feature screening without strong distributional assumptions, thus enhancing the method’s generalizability and robustness. In contrast, the conditional regression error difference 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is dependent on the data distribution. As such, when 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <msub> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math>, the statistic becomes less sensitive to distributional shifts, since the distribution-free term 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> dominates. By adjusting the weights 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math>, users can flexibly balance robustness and model interpretability, making the statistic adaptable to different data structures and application needs across a wide range of scenarios.</p>
   </sec>
   <sec id="s3_2">
    <title>3.2. Feature Screening</title>
    <p>In the following sections, we establish the theoretical properties of the proposed dynamic conditional feature screening (DCFS) procedure. Prior studies, including those by Fan and Lv <xref ref-type="bibr" rid="scirp.142110-1">
      [1]
     </xref> and Ni and Fang <xref ref-type="bibr" rid="scirp.142110-5">
      [5]
     </xref>, have demonstrated that the sure screening property plays a central role in validating the effectiveness of independent screening methods. Therefore, it is essential to rigorously justify the theoretical reliability of the DCFS method. To this end, we introduce a set of regularity conditions under which the screening performance of DCFS can be formally guaranteed. While these conditions may not be the weakest possible, they are primarily imposed to facilitate the technical derivation and proof of the theoretical results.</p>
    <p>Assuming the following conditions:</p>
    <p>The joint density of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
        Z 
      </mtext> 
     </math> is continuous and bounded, i.e.,</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mtext>
          Z 
        </mtext> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mtext>
          z 
        </mtext> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         C 
       </mi> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ‖ 
        </mo> 
        <mrow> 
         <mo>
           ∇ 
         </mo> 
         <msub> 
          <mi>
            f 
          </mi> 
          <mtext>
            Z 
          </mtext> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mtext>
            z 
          </mtext> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ‖ 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         C 
       </mi> 
      </mrow> 
     </math>. (22)</p>
    <p>This ensures the stability of the conditional variable distribution and avoids failure of density estimation in high-dimensional settings.</p>
    <p>The estimation errors satisfy</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ‖ 
        </mo> 
        <mrow> 
         <mover accent="true"> 
          <mi>
            I 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
         <mo>
           − 
         </mo> 
         <mi>
           I 
         </mi> 
        </mrow> 
        <mo>
          ‖ 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mrow> 
        <mo>
          ‖ 
        </mo> 
        <mrow> 
         <mi>
           Δ 
         </mi> 
         <mover accent="true"> 
          <mi>
            E 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
         <mo>
           − 
         </mo> 
         <mi>
           Δ 
         </mi> 
         <mi>
           E 
         </mi> 
        </mrow> 
        <mo>
          ‖ 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi mathvariant="script">
          O 
        </mi> 
        <mi>
          p 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msup> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <mi>
             γ 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (23)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         γ 
       </mi> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>. This condition requires the estimation errors of deep learning models (VAE and MLP) to decay at an exponential rate with increasing sample size, ensuring the convergence of the statistic.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mi>
           min 
         </mi> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         2 
       </mn> 
       <mi>
         c 
       </mi> 
       <msup> 
        <mi>
          n 
        </mi> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           τ 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (24)</p>
    <p>The unified statistics of active variables must be significantly larger than the noise level, preventing them from being masked by high-dimensional noise.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mi>
           sup 
         </mi> 
        </mrow> 
        <mi>
          j 
        </mi> 
       </munder> 
       <mi mathvariant="double-struck">
         E 
       </mi> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mi>
           exp 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             λ 
           </mi> 
           <mrow> 
            <mo>
              ‖ 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                X 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ‖ 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
       <mo>
         &lt; 
       </mo> 
       <mi>
         ∞ 
       </mi> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mi>
         λ 
       </mi> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0. 
       </mn> 
      </mrow> 
     </math> (25)</p>
    <p>This controls the influence of outliers and ensures that concentration inequalities for the statistics hold.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mfrac> 
        <mrow> 
         <msub> 
          <mi>
            c 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
        </mrow> 
        <mi>
          R 
        </mi> 
       </mfrac> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         ℙ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           = 
         </mo> 
         <mi>
           r 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mfrac> 
        <mrow> 
         <msub> 
          <mi>
            c 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
        </mrow> 
        <mi>
          R 
        </mi> 
       </mfrac> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         R 
       </mi> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (26)</p>
    <p>This condition avoids class imbalance, which could bias the estimation of dependence and impair fairness in both MIC and regression error difference measures.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0. 
       </mn> 
      </mrow> 
     </math> (27)</p>
    <p>Ensures that conditional mutual information is well-defined and avoids numerical instability caused by zero-probability events.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mn>
          4 
        </mn> 
       </msub> 
       <msup> 
        <mi>
          n 
        </mi> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           ρ 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          x 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mtext>
           
       </mtext> 
       <mtext>
         is 
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         continuous 
       </mtext> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (28)</p>
    <p>It prevents the failure of density estimation under sparse data scenarios and ensures theoretical convergence for methods such as kNN and VAE.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mi>
           lim 
         </mi> 
         <mi>
           inf 
         </mi> 
        </mrow> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           → 
         </mo> 
         <mi>
           ∞ 
         </mi> 
        </mrow> 
       </munder> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <munder> 
          <mrow> 
           <mi>
             min 
           </mi> 
          </mrow> 
          <mrow> 
           <mi>
             j 
           </mi> 
           <mo>
             ∈ 
           </mo> 
           <mi>
             S 
           </mi> 
          </mrow> 
         </munder> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <munder> 
          <mrow> 
           <mi>
             max 
           </mi> 
          </mrow> 
          <mrow> 
           <mi>
             j 
           </mi> 
           <mo>
             ∉ 
           </mo> 
           <mi>
             S 
           </mi> 
          </mrow> 
         </munder> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mi>
         δ 
       </mi> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0. 
       </mn> 
      </mrow> 
     </math> (29)</p>
    <p>Guarantees that the statistics of important variables can be asymptotically separated from those of irrelevant variables, reducing misclassification risk.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         Δ 
       </mi> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <msub> 
        <mi>
          c 
        </mi> 
        <mn>
          5 
        </mn> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0. 
       </mn> 
      </mrow> 
     </math> (30)</p>
    <p>Prevents division by zero during the computation of weights, ensuring the unified statistic is well-defined.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ‖ 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            w 
          </mi> 
          <mn>
            1 
          </mn> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             I 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             Δ 
           </mi> 
           <mi>
             E 
           </mi> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           − 
         </mo> 
         <msubsup> 
          <mi>
            w 
          </mi> 
          <mn>
            1 
          </mn> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msup> 
            <mi>
              I 
            </mi> 
            <mo>
              ′ 
            </mo> 
           </msup> 
           <mo>
             , 
           </mo> 
           <mi>
             Δ 
           </mi> 
           <msup> 
            <mi>
              E 
            </mi> 
            <mo>
              ′ 
            </mo> 
           </msup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ‖ 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         L 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            ‖ 
          </mo> 
          <mrow> 
           <mi>
             I 
           </mi> 
           <mo>
             − 
           </mo> 
           <msup> 
            <mi>
              I 
            </mi> 
            <mo>
              ′ 
            </mo> 
           </msup> 
          </mrow> 
          <mo>
            ‖ 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mrow> 
          <mo>
            ‖ 
          </mo> 
          <mrow> 
           <mi>
             Δ 
           </mi> 
           <mi>
             E 
           </mi> 
           <mo>
             − 
           </mo> 
           <mi>
             Δ 
           </mi> 
           <msup> 
            <mi>
              E 
            </mi> 
            <mo>
              ′ 
            </mo> 
           </msup> 
          </mrow> 
          <mo>
            ‖ 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (31)</p>
    <p>Ensures robustness of the weight function against small estimation errors, avoiding instability in the unified statistic due to minor fluctuations.</p>
    <p>Under the above conditions, we can rigorously establish the reliable screening performance of the DCFS procedure. The detailed proof is presented in the following subsection.</p>
    <p>The sure screening property refers to the asymptotic guarantee that, in high-dimensional settings, all truly important variables (i.e., variables conditionally associated with the response 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>) are retained in the selected feature set with probability tending to one. Below we present the formal proof under the regularity conditions (C1)-(C4), (C9), and (C10).</p>
    <p>Step 1: Error Decomposition and Weight Stability</p>
    <p>Let the true dynamic statistic be denoted as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mi>
         I 
       </mi> 
       <mo>
         + 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (32)</p>
    <p>and its estimator as</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mover accent="true"> 
         <mi>
           w 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mover accent="true"> 
        <mi>
          I 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         + 
       </mo> 
       <msubsup> 
        <mover accent="true"> 
         <mi>
           w 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mtext>
         Δ 
       </mtext> 
       <mover accent="true"> 
        <mi>
          E 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (33)</p>
    <p>Then, the absolute estimation error can be decomposed as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             T 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          | 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <munder> 
        <munder> 
         <mrow> 
          <mrow> 
           <mo>
             | 
           </mo> 
           <mrow> 
            <msubsup> 
             <mover accent="true"> 
              <mi>
                w 
              </mi> 
              <mo>
                ^ 
              </mo> 
             </mover> 
             <mn>
               1 
             </mn> 
             <mrow> 
              <mrow> 
               <mo>
                 ( 
               </mo> 
               <mi>
                 j 
               </mi> 
               <mo>
                 ) 
               </mo> 
              </mrow> 
             </mrow> 
            </msubsup> 
            <mo>
              − 
            </mo> 
            <msubsup> 
             <mi>
               w 
             </mi> 
             <mn>
               1 
             </mn> 
             <mrow> 
              <mrow> 
               <mo>
                 ( 
               </mo> 
               <mi>
                 j 
               </mi> 
               <mo>
                 ) 
               </mo> 
              </mrow> 
             </mrow> 
            </msubsup> 
           </mrow> 
           <mo>
             | 
           </mo> 
          </mrow> 
          <mo>
            ⋅ 
          </mo> 
          <mi>
            I 
          </mi> 
         </mrow> 
         <mo stretchy="true">
           ︸ 
         </mo> 
        </munder> 
        <mrow> 
         <mtext>
           weight error 
         </mtext> 
        </mrow> 
       </munder> 
       <mo>
         + 
       </mo> 
       <munder> 
        <munder> 
         <mrow> 
          <mrow> 
           <mo>
             | 
           </mo> 
           <mrow> 
            <msubsup> 
             <mover accent="true"> 
              <mi>
                w 
              </mi> 
              <mo>
                ^ 
              </mo> 
             </mover> 
             <mn>
               2 
             </mn> 
             <mrow> 
              <mrow> 
               <mo>
                 ( 
               </mo> 
               <mi>
                 j 
               </mi> 
               <mo>
                 ) 
               </mo> 
              </mrow> 
             </mrow> 
            </msubsup> 
            <mo>
              − 
            </mo> 
            <msubsup> 
             <mi>
               w 
             </mi> 
             <mn>
               2 
             </mn> 
             <mrow> 
              <mrow> 
               <mo>
                 ( 
               </mo> 
               <mi>
                 j 
               </mi> 
               <mo>
                 ) 
               </mo> 
              </mrow> 
             </mrow> 
            </msubsup> 
           </mrow> 
           <mo>
             | 
           </mo> 
          </mrow> 
          <mo>
            ⋅ 
          </mo> 
          <mtext>
            Δ 
          </mtext> 
          <mi>
            E 
          </mi> 
         </mrow> 
         <mo stretchy="true">
           ︸ 
         </mo> 
        </munder> 
        <mrow> 
         <mtext>
           weight error 
         </mtext> 
        </mrow> 
       </munder> 
       <mo>
         + 
       </mo> 
       <munder> 
        <munder> 
         <mrow> 
          <msubsup> 
           <mi>
             w 
           </mi> 
           <mn>
             1 
           </mn> 
           <mrow> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               j 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
          </msubsup> 
          <mrow> 
           <mo>
             | 
           </mo> 
           <mrow> 
            <mover accent="true"> 
             <mi>
               I 
             </mi> 
             <mo>
               ^ 
             </mo> 
            </mover> 
            <mo>
              − 
            </mo> 
            <mi>
              I 
            </mi> 
           </mrow> 
           <mo>
             | 
           </mo> 
          </mrow> 
         </mrow> 
         <mo stretchy="true">
           ︸ 
         </mo> 
        </munder> 
        <mrow> 
         <mtext>
           MI estimation error 
         </mtext> 
        </mrow> 
       </munder> 
       <mo>
         + 
       </mo> 
       <munder> 
        <munder> 
         <mrow> 
          <msubsup> 
           <mi>
             w 
           </mi> 
           <mn>
             2 
           </mn> 
           <mrow> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               j 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
           </mrow> 
          </msubsup> 
          <mrow> 
           <mo>
             | 
           </mo> 
           <mrow> 
            <mtext>
              Δ 
            </mtext> 
            <mover accent="true"> 
             <mi>
               E 
             </mi> 
             <mo>
               ^ 
             </mo> 
            </mover> 
            <mo>
              − 
            </mo> 
            <mtext>
              Δ 
            </mtext> 
            <mi>
              E 
            </mi> 
           </mrow> 
           <mo>
             | 
           </mo> 
          </mrow> 
         </mrow> 
         <mo stretchy="true">
           ︸ 
         </mo> 
        </munder> 
        <mrow> 
         <mtext>
           regression error 
         </mtext> 
        </mrow> 
       </munder> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (34)</p>
    <p>From conditions (C9) and (C10), the weight estimation error is Lipschitz continuous and bounded:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mrow> 
         <msubsup> 
          <mover accent="true"> 
           <mi>
             w 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mi>
            k 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mo>
           − 
         </mo> 
         <msubsup> 
          <mi>
            w 
          </mi> 
          <mi>
            k 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
        </mrow> 
        <mo>
          | 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mfrac> 
        <mi>
          L 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            c 
          </mi> 
          <mn>
            5 
          </mn> 
         </msub> 
        </mrow> 
       </mfrac> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              I 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             − 
           </mo> 
           <mi>
             I 
           </mi> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <mtext>
             Δ 
           </mtext> 
           <mover accent="true"> 
            <mi>
              E 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             − 
           </mo> 
           <mtext>
             Δ 
           </mtext> 
           <mi>
             E 
           </mi> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         for 
       </mtext> 
       <mtext>
           
       </mtext> 
       <mi>
         k 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         , 
       </mo> 
       <mn>
         2. 
       </mn> 
      </mrow> 
     </math> (34)</p>
    <p>Step 2: Key Lemmas - Concentration Inequalities</p>
    <p>We invoke the following lemmas to control the stochastic error terms:</p>
    <p>Under conditions (C1)–(C2), there exists a constant 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mi>
          I 
        </mi> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> such that</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ℙ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              I 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             − 
           </mo> 
           <mi>
             I 
           </mi> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
         <mo>
           ≥ 
         </mo> 
         <mi>
           ϵ 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mn>
         2 
       </mn> 
       <mi>
         exp 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mi>
            I 
          </mi> 
         </msub> 
         <mi>
           n 
         </mi> 
         <msup> 
          <mi>
            ϵ 
          </mi> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (35)</p>
    <p>Under conditions (C1)-(C4), there exists a constant 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mi>
          E 
        </mi> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> such that</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ℙ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <mtext>
             Δ 
           </mtext> 
           <mover accent="true"> 
            <mi>
              E 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             − 
           </mo> 
           <mtext>
             Δ 
           </mtext> 
           <mi>
             E 
           </mi> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
         <mo>
           ≥ 
         </mo> 
         <mi>
           ϵ 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mn>
         2 
       </mn> 
       <mi>
         exp 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mi>
            E 
          </mi> 
         </msub> 
         <mi>
           n 
         </mi> 
         <msup> 
          <mi>
            ϵ 
          </mi> 
          <mn>
            2 
          </mn> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (36)</p>
    <p>Proof Techniques:</p>
    <p>Step 3: Uniform Error Bound</p>
    <p>Combining the four error terms, we derive the total estimation bound:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"> <mrow> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             T 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          | 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mfrac> 
          <mrow> 
           <mn>
             2 
           </mn> 
           <mi>
             L 
           </mi> 
          </mrow> 
          <mrow> 
           <msub> 
            <mi>
              c 
            </mi> 
            <mn>
              5 
            </mn> 
           </msub> 
          </mrow> 
         </mfrac> 
         <mo>
           + 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              I 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             − 
           </mo> 
           <mi>
             I 
           </mi> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <mtext>
             Δ 
           </mtext> 
           <mover accent="true"> 
            <mi>
              E 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             − 
           </mo> 
           <mtext>
             Δ 
           </mtext> 
           <mi>
             E 
           </mi> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (37)</p>
    <p>Set 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         c 
       </mi> 
       <msup> 
        <mi>
          n 
        </mi> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           τ 
         </mi> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> for some 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         τ 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mfrac> 
          <mn>
            1 
          </mn> 
          <mn>
            2 
          </mn> 
         </mfrac> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. Then, using union bounds over Lemmas 3.1</p>
    <p>and 3.2:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ℙ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <msub> 
            <mover accent="true"> 
             <mi>
               T 
             </mi> 
             <mo>
               ^ 
             </mo> 
            </mover> 
            <mi>
              j 
            </mi> 
           </msub> 
           <mo>
             − 
           </mo> 
           <msub> 
            <mi>
              T 
            </mi> 
            <mi>
              j 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
         <mo>
           ≥ 
         </mo> 
         <mi>
           ϵ 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≤ 
       </mo> 
       <mn>
         4 
       </mn> 
       <mi>
         exp 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
         <msup> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mn>
             1 
           </mn> 
           <mo>
             − 
           </mo> 
           <mn>
             2 
           </mn> 
           <mi>
             τ 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (38)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mtext>
         min 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mi>
            I 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mi>
            E 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ⋅ 
       </mo> 
       <msup> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mn>
             2 
           </mn> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mrow> 
              <mrow> 
               <mn>
                 2 
               </mn> 
               <mi>
                 L 
               </mi> 
              </mrow> 
              <mo>
                / 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  c 
                </mi> 
                <mn>
                  5 
                </mn> 
               </msub> 
              </mrow> 
             </mrow> 
             <mo>
               + 
             </mo> 
             <mn>
               1 
             </mn> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mn>
           2 
         </mn> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>.</p>
    <p>Step 4: Maximal Deviation Over All Features</p>
    <p>Apply the union bound over all 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math> variables:</p>
    <p><img width="281.25" src="https://html.scirp.org/file/1241932-rId336.svg?20250610043620"> (39)</img></p>
    <p>By Condition (C3), we assume</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mtext>
           min 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         2 
       </mn> 
       <mi>
         c 
       </mi> 
       <msup> 
        <mi>
          n 
        </mi> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           τ 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (40)</p>
    <p>Thus, for all active variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         j 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mi>
         S 
       </mi> 
      </mrow> 
     </math>, their estimators satisfy</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≥ 
       </mo> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         − 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ≥ 
       </mo> 
       <mi>
         c 
       </mi> 
       <msup> 
        <mi>
          n 
        </mi> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           τ 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (41)</p>
    <p>For inactive variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         j 
       </mi> 
       <mo>
         ∉ 
       </mo> 
       <mi>
         S 
       </mi> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>, and</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≤ 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         c 
       </mi> 
       <msup> 
        <mi>
          n 
        </mi> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mi>
           τ 
         </mi> 
        </mrow> 
       </msup> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (42)</p>
    <p>Define the selection rule:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mi>
          S 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           : 
         </mo> 
         <msub> 
          <mover accent="true"> 
           <mi>
             T 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           &gt; 
         </mo> 
         <mi>
           c 
         </mi> 
         <msup> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <mi>
             τ 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (43)</p>
    <p>Then the screening rule guarantees that all important features are selected:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ℙ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           S 
         </mi> 
         <mo>
           ⊆ 
         </mo> 
         <mover accent="true"> 
          <mi>
            S 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <mn>
         4 
       </mn> 
       <msub> 
        <mi>
          s 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
       <mi>
         exp 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
         <msup> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mn>
             1 
           </mn> 
           <mo>
             − 
           </mo> 
           <mn>
             2 
           </mn> 
           <mi>
             τ 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (44)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          s 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mi>
          S 
        </mi> 
        <mo>
          | 
        </mo> 
       </mrow> 
      </mrow> 
     </math> is the number of active features.</p>
    <p>Let 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mrow> 
         <msub> 
          <mi>
            C 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
        </mrow> 
        <mo>
          / 
        </mo> 
        <mrow> 
         <mi>
           log 
         </mi> 
         <mn>
           2 
         </mn> 
        </mrow> 
       </mrow> 
      </mrow> 
     </math>, and we conclude:</p>
    <p>Theorem 1 (Sure Screening Property)</p>
    <p>Under Conditions (C1)-(C4), (C9), and (C10), the dynamic statistic</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mtext>
           dynamic 
         </mtext> 
        </mrow> 
       </msubsup> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <msubsup> 
        <mi>
          w 
        </mi> 
        <mn>
          2 
        </mn> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (45)</p>
    <p>satisfies the following probabilistic bound:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ℙ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           S 
         </mi> 
         <mo>
           ⊆ 
         </mo> 
         <mover accent="true"> 
          <mi>
            S 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <mi mathvariant="script">
         O 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            s 
          </mi> 
          <mi>
            n 
          </mi> 
         </msub> 
         <mi>
           exp 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <msub> 
            <mi>
              C 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
           <msup> 
            <mi>
              n 
            </mi> 
            <mrow> 
             <mn>
               1 
             </mn> 
             <mo>
               − 
             </mo> 
             <mn>
               2 
             </mn> 
             <mi>
               τ 
             </mi> 
            </mrow> 
           </msup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (46)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        S 
      </mi> 
     </math> is the true active set, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         S 
       </mi> 
       <mo>
         ^ 
       </mo> 
      </mover> 
     </math> is the selected feature set, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          s 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mi>
          S 
        </mi> 
        <mo>
          | 
        </mo> 
       </mrow> 
      </mrow> 
     </math>,</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         τ 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mfrac> 
          <mn>
            1 
          </mn> 
          <mn>
            2 
          </mn> 
         </mfrac> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          C 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> is a constant.</p>
    <p>This inequality implies that the probability of missing any important variable decays exponentially as sample size 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        n 
      </mi> 
     </math> increases, provided the signal strength is not too weak. Additionally, the dimensionality 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math> is allowed to grow at an exponential rate, i.e., 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         p 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="script">
         O 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mtext>
           exp 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msup> 
            <mi>
              n 
            </mi> 
            <mrow> 
             <mn>
               1 
             </mn> 
             <mo>
               − 
             </mo> 
             <mn>
               2 
             </mn> 
             <mi>
               τ 
             </mi> 
            </mrow> 
           </msup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, which demonstrates the scalability and robustness of the proposed screening method in ultra-high-dimensional regimes.</p>
    <p>The ranking consistency property states that, as the sample size increases, the estimated importance scores 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> for the truly important variables remain consistently larger than those for the unimportant ones. We provide a formal proof under Conditions (C5)-(C8) and the dynamic weighting conditions (C9)-(C10).</p>
    <p>Step 1: Strong Consistency of the Statistic</p>
    <p>By Conditions (C5)-(C7) and Lemma 3.3 (Strong consistency of density estimators), we have:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           f 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mrow> 
         <mtext>
           a 
         </mtext> 
         <mtext>
           .s 
         </mtext> 
         <mtext>
           . 
         </mtext> 
        </mrow> 
       </mover> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <msub> 
        <mover accent="true"> 
         <mi>
           f 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           ∣ 
         </mo> 
         <mi>
           y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mrow> 
         <mtext>
           a 
         </mtext> 
         <mtext>
           .s 
         </mtext> 
         <mtext>
           . 
         </mtext> 
        </mrow> 
       </mover> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           | 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           x 
         </mi> 
         <mo>
           | 
         </mo> 
         <mi>
           y 
         </mi> 
         <mo>
           , 
         </mo> 
         <mtext>
           z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (47)</p>
    <p>As a result, the conditional mutual information estimator converges almost surely:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mi>
          I 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mrow> 
         <mtext>
           a 
         </mtext> 
         <mtext>
           .s 
         </mtext> 
         <mtext>
           . 
         </mtext> 
        </mrow> 
       </mover> 
       <mi>
         I 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ; 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (48)</p>
    <p>and the regression error difference estimator satisfies:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mover accent="true"> 
        <mi>
          E 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mrow> 
         <mtext>
           a 
         </mtext> 
         <mtext>
           .s 
         </mtext> 
         <mtext>
           . 
         </mtext> 
        </mrow> 
       </mover> 
       <mtext>
         Δ 
       </mtext> 
       <mi>
         E 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
         <mo>
           | 
         </mo> 
         <mtext>
           Z 
         </mtext> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (49)</p>
    <p>Since the dynamic weights are Lipschitz continuous (Condition C10), applying the continuous mapping theorem, we obtain:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mover accent="true"> 
         <mi>
           I 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mrow> 
         <mover accent="true"> 
          <mi>
            I 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
         <mo>
           + 
         </mo> 
         <mtext>
           Δ 
         </mtext> 
         <mover accent="true"> 
          <mi>
            E 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
        </mrow> 
       </mfrac> 
       <mo>
         ⋅ 
       </mo> 
       <mover accent="true"> 
        <mi>
          I 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         + 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Δ 
         </mtext> 
         <mover accent="true"> 
          <mi>
            E 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
        </mrow> 
        <mrow> 
         <mover accent="true"> 
          <mi>
            I 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
         <mo>
           + 
         </mo> 
         <mtext>
           Δ 
         </mtext> 
         <mover accent="true"> 
          <mi>
            E 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
        </mrow> 
       </mfrac> 
       <mo>
         ⋅ 
       </mo> 
       <mtext>
         Δ 
       </mtext> 
       <mover accent="true"> 
        <mi>
          E 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mover> 
        <mo>
          → 
        </mo> 
        <mrow> 
         <mtext>
           a 
         </mtext> 
         <mtext>
           .s 
         </mtext> 
         <mtext>
           . 
         </mtext> 
        </mrow> 
       </mover> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (50)</p>
    <p>Step 2: Stability of Signal Separation</p>
    <p>From Condition (C8), there exists 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         δ 
       </mi> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          N 
        </mi> 
        <mn>
          0 
        </mn> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math> such that for all 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         &gt; 
       </mo> 
       <msub> 
        <mi>
          N 
        </mi> 
        <mn>
          0 
        </mn> 
       </msub> 
      </mrow> 
     </math>:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mtext>
           min 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         − 
       </mo> 
       <munder> 
        <mrow> 
         <mtext>
           max 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∉ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≥ 
       </mo> 
       <mi>
         δ 
       </mi> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (51)</p>
    <p>For any 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mrow> 
          <mi>
            δ 
          </mi> 
          <mo>
            / 
          </mo> 
          <mn>
            4 
          </mn> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, the strong consistency of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> implies that there exists 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          N 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <msub> 
        <mi>
          N 
        </mi> 
        <mn>
          0 
        </mn> 
       </msub> 
      </mrow> 
     </math> such that for all 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         &gt; 
       </mo> 
       <msub> 
        <mi>
          N 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math>:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             T 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          | 
        </mo> 
       </mrow> 
       <mo>
         &lt; 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         a 
       </mtext> 
       <mtext>
         .s 
       </mtext> 
       <mtext>
         . 
       </mtext> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mo>
         ∀ 
       </mo> 
       <mi>
         j 
       </mi> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (52)</p>
    <p>Therefore, for active variables:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mtext>
           min 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≥ 
       </mo> 
       <munder> 
        <mrow> 
         <mtext>
           min 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         − 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ≥ 
       </mo> 
       <munder> 
        <mrow> 
         <mtext>
           max 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∉ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         + 
       </mo> 
       <mi>
         δ 
       </mi> 
       <mo>
         − 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (53)</p>
    <p>and for inactive variables:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mtext>
           max 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∉ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≤ 
       </mo> 
       <munder> 
        <mrow> 
         <mtext>
           max 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∉ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         + 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (54)</p>
    <p>Setting 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mi>
          δ 
        </mi> 
        <mo>
          / 
        </mo> 
        <mn>
          4 
        </mn> 
       </mrow> 
      </mrow> 
     </math>, we obtain:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mtext>
           min 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ≥ 
       </mo> 
       <mfrac> 
        <mrow> 
         <mn>
           3 
         </mn> 
         <mi>
           δ 
         </mi> 
        </mrow> 
        <mn>
          4 
        </mn> 
       </mfrac> 
       <mo>
         &gt; 
       </mo> 
       <mfrac> 
        <mi>
          δ 
        </mi> 
        <mn>
          4 
        </mn> 
       </mfrac> 
       <mo>
         ≥ 
       </mo> 
       <munder> 
        <mrow> 
         <mtext>
           max 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∉ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         , 
       </mo> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         a 
       </mtext> 
       <mtext>
         .s 
       </mtext> 
       <mtext>
         . 
       </mtext> 
      </mrow> 
     </math> (55)</p>
    <p>Step 3: Borel-Cantelli Lemma</p>
    <p>Since</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle displaystyle="true"> 
        <msubsup> 
         <mo>
           ∑ 
         </mo> 
         <mrow> 
          <mi>
            n 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           ∞ 
         </mi> 
        </msubsup> 
        <mrow> 
         <mi>
           ℙ 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mrow> 
            <mo>
              | 
            </mo> 
            <mrow> 
             <msub> 
              <mover accent="true"> 
               <mi>
                 T 
               </mi> 
               <mo>
                 ^ 
               </mo> 
              </mover> 
              <mi>
                j 
              </mi> 
             </msub> 
             <mo>
               − 
             </mo> 
             <msub> 
              <mi>
                T 
              </mi> 
              <mi>
                j 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              | 
            </mo> 
           </mrow> 
           <mo>
             ≥ 
           </mo> 
           <mi>
             ϵ 
           </mi> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </mstyle> 
       <mo>
         ≤ 
       </mo> 
       <mstyle displaystyle="true"> 
        <msubsup> 
         <mo>
           ∑ 
         </mo> 
         <mrow> 
          <mi>
            n 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           ∞ 
         </mi> 
        </msubsup> 
        <mrow> 
         <mn>
           4 
         </mn> 
         <mi>
           exp 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mo>
             − 
           </mo> 
           <msub> 
            <mi>
              C 
            </mi> 
            <mn>
              2 
            </mn> 
           </msub> 
           <msup> 
            <mi>
              n 
            </mi> 
            <mrow> 
             <mn>
               1 
             </mn> 
             <mo>
               − 
             </mo> 
             <mn>
               2 
             </mn> 
             <mi>
               τ 
             </mi> 
            </mrow> 
           </msup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </mstyle> 
       <mo>
         &lt; 
       </mo> 
       <mi>
         ∞ 
       </mi> 
       <mo>
         , 
       </mo> 
      </mrow> 
     </math> (56)</p>
    <p>the Borel-Cantelli lemma implies that the event 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          | 
        </mo> 
        <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             T 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           − 
         </mo> 
         <msub> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          | 
        </mo> 
       </mrow> 
       <mo>
         ≥ 
       </mo> 
       <mi>
         ϵ 
       </mi> 
      </mrow> 
     </math> only occurs finitely often. Hence:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mi>
           lim 
         </mi> 
        </mrow> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           → 
         </mo> 
         <mi>
           ∞ 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         a 
       </mtext> 
       <mtext>
         .s 
       </mtext> 
       <mtext>
         . 
       </mtext> 
      </mrow> 
     </math> (57)</p>
    <p>Theorem 2 (Ranking Consistency)</p>
    <p>Under Conditions (C5)-(C8) and the dynamic weighting conditions (C9)-(C10), as 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         → 
       </mo> 
       <mi>
         ∞ 
       </mi> 
      </mrow> 
     </math>, we have:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <munder> 
        <mrow> 
         <mtext>
           min 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         &gt; 
       </mo> 
       <munder> 
        <mrow> 
         <mtext>
           max 
         </mtext> 
        </mrow> 
        <mrow> 
         <mi>
           j 
         </mi> 
         <mo>
           ∉ 
         </mo> 
         <mi>
           S 
         </mi> 
        </mrow> 
       </munder> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         almost 
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         surely 
       </mtext> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (58)</p>
    <p>This result shows that the estimated statistics 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> converge almost surely to the true values 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          T 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math>, and due to the Lipschitz continuity of the weight function (C10), the convergence of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         I 
       </mi> 
       <mo>
         ^ 
       </mo> 
      </mover> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mover accent="true"> 
        <mi>
          E 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
      </mrow> 
     </math> is stably transferred to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mover accent="true"> 
         <mi>
           T 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math>. With a guaranteed signal separation (C8), the estimation errors cannot disrupt the correct variable ordering. Thus, the proposed screening method consistently ranks important variables above unimportant ones. This ranking stability ensures that the selection outcome remains reliable across varying sample realizations and is not overly sensitive to small fluctuations in the data, further enhancing the robustness and interpretability of the feature screening procedure.</p>
   </sec>
   <sec id="s3_3">
    <title>3.3. Practical Justification of Theoretical Assumptions</title>
    <p>Sections 3.1 and 3.2 have established the theoretical foundation of the proposed DCFS method, including its sure screening and ranking consistency properties. These results are derived under a set of ten technical assumptions, denoted as Conditions (C1) through (C10). While these conditions facilitate rigorous theoretical analysis, their practical plausibility is essential for the method’s real-world applicability. In this section, we examine the feasibility of each assumption in empirical settings and offer practical guidelines for their verification.</p>
    <p>(1) Discussion of Individual Assumptions</p>
    <p>This condition is typically satisfied in most real-world datasets in economics, finance, and biomedicine, where variables often follow approximately continuous distributions. Occasional extreme values can be effectively handled through preprocessing techniques such as normalization, robust transformations, or winsorization.</p>
    <p>Although theoretically strong, this assumption is often met or closely approximated in practice when using modern deep neural networks with appropriate architectures and optimizers (e.g., Adam, RMSProp). Empirical convergence behavior can be validated through loss curve diagnostics and cross-validation <xref ref-type="bibr" rid="scirp.142110-17">
      [17]
     </xref>.</p>
    <p>Weak-signal variables may be dominated by noise in high-dimensional settings. In practice, preliminary filtering using correlation screening or statistical significance testing can help satisfy this assumption by discarding irrelevant features before applying DCFS.</p>
    <p>This condition can be ensured through standard data preprocessing methods such as outlier truncation or robust scaling, especially when dealing with heavy-tailed distributions commonly observed in high-dimensional data.</p>
    <p>While this condition is not relevant for regression problems, it can be addressed in classification tasks using sampling strategies (e.g., SMOTE, undersampling) or by introducing class weights into the loss function.</p>
    <p>In large-sample scenarios, these assumptions are generally satisfied or approximated, especially when the data are reasonably well-distributed. Visual inspection using kernel density plots or low-dimensional projections can assist in assessing these conditions.</p>
    <p>This assumption is more restrictive, as real data may not always exhibit a clear margin between important and irrelevant features. In Section 4, we conduct simulation studies to evaluate the robustness of DCFS under mild violations of this condition.</p>
    <p>These assumptions are easy to enforce through regularization strategies during implementation (e.g., bounding gradients, avoiding near-zero denominators), and typically pose no obstacle in practice.</p>
    <p>(2) Practical Guidelines for Assumption Verification</p>
    <p>To assess whether a dataset satisfies the theoretical assumptions required by DCFS, we recommend the following practical steps:</p>
    <p>Use histograms, boxplots, and kernel density estimates to check distribution continuity, detect outliers, and assess marginal and conditional density behavior.</p>
    <p>Perform simple regression or correlation analysis to identify features with weak or negligible association with the response variable.</p>
    <p>During training of the VAE and MLP components, monitor the loss curve. A consistently decreasing trajectory (ideally exponential) indicates that the estimation error behaves as required.</p>
    <p>In classification tasks, compute the class proportions and apply balancing techniques if significant imbalance is detected.</p>
    <p>These guidelines provide a practical roadmap for evaluating the applicability of DCFS in empirical contexts, ensuring that the underlying assumptions are met and that the theoretical guarantees translate effectively to real-world performance.</p>
   </sec>
  </sec><sec id="s4">
   <title>4. Numerical Simulation Experiment</title>
   <p>This chapter is based on the Dynamic Conditional Feature Selection (DCFS) method proposed in this paper. Through numerical simulation experiments, the effectiveness and advantages of DCFS in identifying important variables are verified. The experiment designed three different simulation data scenarios: linear model, nonlinear model, and mixed linear and nonlinear model. DCFS was compared and analyzed with existing classical feature selection methods under multiple evaluation indicators. The experimental results clearly demonstrated the stability and advantages of the proposed method.</p>
   <sec id="s4_1">
    <title>4.1. Scenario 1: Linear Model</title>
    <p>In the linear model scenario, we generate simulation data based on the following model:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         Y 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         X 
       </mi> 
       <mi>
         β 
       </mi> 
       <mo>
         + 
       </mo> 
       <mi>
         Z 
       </mi> 
       <mi>
         γ 
       </mi> 
       <mo>
         + 
       </mo> 
       <mi>
         ϵ 
       </mi> 
      </mrow> 
     </math> (59)</p>
    <p>Among them:</p>
    <p>We set some of the predictor variables (X) as active variables (variables that truly affect Y), and the other variables as pure noise. The response variable (Y) is generated by a linear combination of these active variables and all conditional variables (Z).</p>
    <p>Three different sample sizes were selected for simulation experiments for comparison, with the following specific settings: sample size: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mn>
           200 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           300 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           500 
         </mn> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math>; Prediction variable dimension: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         p 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1000 
       </mn> 
      </mrow> 
     </math>, where the first 10 variables are active variables and their coefficients are 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          β 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         ~ 
       </mo> 
       <mi>
         U 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mo>
           − 
         </mo> 
         <mn>
           5 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           5 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>; Dimension of conditional variable: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         q 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         5 
       </mn> 
      </mrow> 
     </math>, set 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         γ 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>; Data generation method: X and Z elements are independently and identically distributed in 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, noise, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ~ 
       </mo> 
       <mi>
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>. The experiment was independently repeated 100 times for each sample size, and the average of all outcome indicators was taken to reduce the impact of random fluctuations on the experimental results.</p>
    <p>In this linear scenario, select the following feature filtering methods for performance comparison with DCFS:</p>
    <p>In this scenario, we evaluate the performance of the screening method using the following three indicators:</p>
    <p>(1) True Positive Rate (TPR): The proportion of correctly identified truly important features to all truly important features. The mathematical definition is:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         TPR 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FN 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> (60)</p>
    <p>TP (True Positive, true example): Refers to the number of important variables selected by the method that are actually significant in the model.</p>
    <p>FN (False Negative): The number of important variables that the method failed to recognize.</p>
    <p>TPR represents the sensitivity of the method. The closer the value is to 1, the more accurate the screening method is in capturing all true signal variables, and the lower the risk of important variables being missed. Ideally, TPR should be close to 1.0, which corresponds to the Sure Screening Property requirement of feature screening methods.</p>
    <p>(2) False Discovery Rate (FDR): FDR describes the proportion of features selected by a filtering method that are actually irrelevant noise variables, and is a measure of the “false alarm” situation in the filtering process. The specific definition of FDR is:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         FDR 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           FP 
         </mtext> 
        </mrow> 
        <mrow> 
         <mtext>
           TP 
         </mtext> 
         <mo>
           + 
         </mo> 
         <mtext>
           FP 
         </mtext> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> (61)</p>
    <p>FP (False Positive): The actual number of irrelevant noise variables selected by the method.</p>
    <p>TP As mentioned above. The lower the FDR, the higher the accuracy of the method’s screening results, which means that more of the selected variables are real signals rather than irrelevant variables.</p>
    <p>Ideally, we would like this ratio to be as low as possible, meaning that most of the selected variables are truly important to the model rather than noise.</p>
    <p>(3) Ranking Consistency (RC): It reflects the stability of feature selection methods against random fluctuations in data. Specifically, it measures the stability of each important feature maintaining a high ranking relative to irrelevant noise features as the sample size increases or during repeated sampling processes. Calculate the reciprocal of the standard deviation of feature ranking in multiple repeated experiments to reflect ranking stability. The specific method is as follows:</p>
    <p>Each feature 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          j 
        </mi> 
       </msub> 
      </mrow> 
     </math> has a ranking position 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          R 
        </mi> 
        <mi>
          j 
        </mi> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            m 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math> in 100 simulation experiments (the m-th simulation experiment), and we calculate the standard deviation of the fluctuation in the ranking of the active variable in multiple experiments, which is then standardized into a stability score:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mrow> 
         <mtext>
           RC 
         </mtext> 
        </mrow> 
        <mi>
          j 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           std 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msubsup> 
            <mi>
              R 
            </mi> 
            <mi>
              j 
            </mi> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mn>
                1 
              </mn> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
           </msubsup> 
           <mo>
             , 
           </mo> 
           <msubsup> 
            <mi>
              R 
            </mi> 
            <mi>
              j 
            </mi> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mn>
                2 
              </mn> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
           </msubsup> 
           <mo>
             , 
           </mo> 
           <mo>
             ⋯ 
           </mo> 
           <mo>
             , 
           </mo> 
           <msubsup> 
            <mi>
              R 
            </mi> 
            <mi>
              j 
            </mi> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <mn>
                 100 
               </mn> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
           </msubsup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mi>
          p 
        </mi> 
       </mfrac> 
      </mrow> 
     </math> (62)</p>
    <p>Furthermore, the overall ranking consistency of active variables can be taken as the average of all active variables:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         RC 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mn>
          1 
        </mn> 
        <mi>
          k 
        </mi> 
       </mfrac> 
       <mstyle displaystyle="true"> 
        <msubsup> 
         <mo>
           ∑ 
         </mo> 
         <mrow> 
          <mi>
            j 
          </mi> 
          <mo>
            = 
          </mo> 
          <mn>
            1 
          </mn> 
         </mrow> 
         <mi>
           k 
         </mi> 
        </msubsup> 
        <mrow> 
         <msub> 
          <mrow> 
           <mtext>
             RC 
           </mtext> 
          </mrow> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
       </mstyle> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        k 
      </mi> 
     </math> represents the number of active variables (63)</p>
    <p>The range of RC values is [0, 1], with higher values indicating that important variables are more stable during the screening process and less susceptible to random fluctuations or accidental noise. In an ideal situation, the higher the RC, the better, indicating that the method has stronger stability in identifying important variables.</p>
    <p>The experimental results are shown in Table 2:</p>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 2. Performance comparison results of various methods in linear model scenarios.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="13.79%"><p style="text-align:center">Sample size (n)</p></td> 
       <td class="custom-bottom-td acenter" width="12.94%"><p style="text-align:center">method</p></td> 
       <td class="custom-bottom-td acenter" width="24.42%"><p style="text-align:center">True positive rate (TPR)</p></td> 
       <td class="custom-bottom-td acenter" width="24.42%"><p style="text-align:center">False Discovery Rate (FDR)</p></td> 
       <td class="custom-bottom-td acenter" width="24.43%"><p style="text-align:center">Sorting Consistency (RC)</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="13.79%"><p style="text-align:center">200</p></td> 
       <td class="custom-top-td acenter" width="12.94%"><p style="text-align:center">DCFS</p></td> 
       <td class="custom-top-td acenter" width="24.42%"><p style="text-align:center">0.92</p></td> 
       <td class="custom-top-td acenter" width="24.42%"><p style="text-align:center">0.06</p></td> 
       <td class="custom-top-td acenter" width="24.43%"><p style="text-align:center">0.90</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="13.79%"><p style="text-align:center">300</p></td> 
       <td class="custom-top-td acenter" width="12.94%"><p style="text-align:center">DCFS</p></td> 
       <td class="custom-top-td acenter" width="24.42%"><p style="text-align:center">0.95</p></td> 
       <td class="custom-top-td acenter" width="24.42%"><p style="text-align:center">0.05</p></td> 
       <td class="custom-top-td acenter" width="24.43%"><p style="text-align:center">0.93</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.98</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.05</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.96</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.90</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.10</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.85</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.93</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.08</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.88</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.96</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.07</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.91</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">SIS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.85</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.18</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.78</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">SIS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.88</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.15</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.82</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">SIS</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.92</p></td> 
       <td class="acenter" width="24.42%"><p style="text-align:center">0.12</p></td> 
       <td class="acenter" width="24.43%"><p style="text-align:center">0.85</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>The result analysis is as follows: DCFS performs the best, TPR is the highest, FDR is the lowest and stably controlled around the set 5%, and the sorting stability is the highest; CSIS performs well after controlling for the influence of conditional variables, but its false discovery rate and ranking stability are slightly inferior to DCFS; SIS performs the worst due to uncontrolled confounding variables, with the highest FDR, lowest TPR, and worst ranking stability. The experimental results validated the advantage of DCFS in identifying linear important variables, demonstrating the effectiveness and superiority of the proposed method in this paper.</p>
   </sec>
   <sec id="s4_2">
    <title>4.2. Scenario 2: Nonlinear Model</title>
    <p>In scenario 2, we further investigate the performance of the proposed dynamic conditional feature selection method (DCFS) in the case of non-linear dependence between variables and response variables. To this end, the following nonlinear model is constructed to generate simulated data:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         Y 
       </mi> 
       <mo>
         = 
       </mo> 
       <mtext>
         sin 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mn>
         0.5 
       </mn> 
       <msubsup> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
        <mn>
          2 
        </mn> 
       </msubsup> 
       <mo>
         + 
       </mo> 
       <mtext>
         log 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              X 
            </mi> 
            <mn>
              3 
            </mn> 
           </msub> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          Z 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         ϵ 
       </mi> 
      </mrow> 
     </math> (64)</p>
    <p>Among them:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          Z 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <mn>
         0.5 
       </mn> 
       <msubsup> 
        <mi>
          Z 
        </mi> 
        <mn>
          2 
        </mn> 
        <mn>
          2 
        </mn> 
       </msubsup> 
      </mrow> 
     </math> (65)</p>
    <p>In the above model, the response variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> exhibits clear nonlinear relationships with the predictor variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math>, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
      </mrow> 
     </math>. Specifically, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math> is related to the response in a periodic fashion, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math> influences 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> through a quadratic nonlinear relationship, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
      </mrow> 
     </math> contributes via a log-transformed nonlinear effect. Under such a complex nonlinear structure—particularly due to the symmetric, even-function nature of the effects of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
      </mrow> 
     </math>—the linear correlation between each predictor and the response is close to zero or negligible.</p>
    <p>As a result, traditional marginal screening methods based on linear correlation are ineffective in identifying these nonlinear but important variables in this setting. This scenario highlights the need for more flexible and adaptive screening procedures capable of capturing both linear and nonlinear dependencies.</p>
    <p>The parameter settings for this scenario are kept consistent with the previous linear case (Scenario 1) to facilitate direct comparison. Specifically: Sample size: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mn>
           200 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           300 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           500 
         </mn> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math>; Number of candidate predictors: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         p 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1000 
       </mn> 
      </mrow> 
     </math>, among which only three variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
      </mrow> 
     </math> are truly important; Number of conditional variables: 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         q 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         5 
       </mn> 
      </mrow> 
     </math>, with 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math> being informative, while the remaining three are noise variables not involved in the response generation; Data generation: All predictors and conditional variables are independently drawn from a standard normal distribution, i.e., 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>; the noise term 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ~ 
       </mo> 
       <mi>
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>; Repetition and evaluation: The simulation is independently repeated 100 times, and the average performance metrics are reported to ensure robust evaluation.</p>
    <p>Due to the significant nonlinear dependencies involved in this scenario, we specifically chose a classical screening method that can capture any nonlinear relationship to compare with the DCFS method proposed in this paper:</p>
    <p>To effectively evaluate the ability of each method to identify nonlinear signals, we adopt the following three evaluation metrics for this scenario:</p>
    <p>This metric measures the proportion of truly important nonlinear variables ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
      </mrow> 
     </math>) that are successfully identified by the screening method. It is calculated as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         NSDR 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mrow> 
         <mtext>
           Number of times nonlinear variables are selected across 100 runs 
         </mtext> 
        </mrow> 
        <mrow> 
         <mn>
           100 
         </mn> 
         <mo>
           × 
         </mo> 
         <mn>
           3 
         </mn> 
        </mrow> 
       </mfrac> 
       <mo>
         . 
       </mo> 
      </mrow> 
     </math> (66)</p>
    <p>The experimental results are shown in <xref ref-type="table" rid="table3">
      Table 3
     </xref>:</p>
    <p>From the above results, it can be seen that the DCFS method performs the best in detecting nonlinear signals, with a significantly higher detection rate than other methods, especially reaching a detection rate of 99% when the sample size increases to 500; In terms of controlling false discovery rate, DCFS stably controls FDR at the set target (about 5%), significantly better than DC-SIS and IG-SIS; In terms of sorting stability, DCFS performs the best with the lowest ranking standard deviation; IG-SIS showed the greatest fluctuation in repeated experiments, while DC-SIS was in the middle but still inferior to DCFS.</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 3. Performance comparison results of various methods in nonlinear model scenarios.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="13.79%"><p style="text-align:center">Sample size (n)</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="12.94%"><p style="text-align:center">method</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="25.86%"><p style="text-align:center">Nonlinear signal detection rate (%)</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="23.70%"><p style="text-align:center">False Discovery Rate (FDR) (%)</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="23.70%"><p style="text-align:center">Sorting Consistency (RC)</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="13.79%"><p style="text-align:center">200</p></td> 
       <td class="custom-top-td acenter" width="12.94%"><p style="text-align:center">DCFS</p></td> 
       <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">95</p></td> 
       <td class="custom-top-td acenter" width="23.70%"><p style="text-align:center">5</p></td> 
       <td class="custom-top-td acenter" width="23.70%"><p style="text-align:center">0.92</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">97</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">5</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.95</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">99</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">5</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.97</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">94</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">20</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.90</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">96</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">18</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.93</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">98</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">15</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.95</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">IG-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">90</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">22</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.87</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="13.79%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.94%"><p style="text-align:center">IG-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">92</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">19</p></td> 
       <td class="acenter" width="23.70%"><p style="text-align:center">0.89</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="13.79%"><p style="text-align:center">500</p></td> 
       <td class="custom-bottom-td acenter" width="12.94%"><p style="text-align:center">IG-SIS</p></td> 
       <td class="custom-bottom-td acenter" width="25.86%"><p style="text-align:center">95</p></td> 
       <td class="custom-bottom-td acenter" width="23.70%"><p style="text-align:center">17</p></td> 
       <td class="custom-bottom-td acenter" width="23.70%"><p style="text-align:center">0.91</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>In summary, the numerical simulation results of the nonlinear model scenario show that the DCFS method proposed in this paper can not only stably and effectively detect important signal variables when nonlinear relationships dominate, but also significantly outperform existing classical nonlinear screening methods with lower false representation rates and higher ranking stability.</p>
   </sec>
   <sec id="s4_3">
    <title>4.3. Scenario 3: Hybrid Model</title>
    <p>To further investigate the performance of various feature selection methods in complex data structures, we construct a hybrid model that combines linear and nonlinear relationships. The specific generation formula is as follows:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         Y 
       </mi> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <msubsup> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
        <mn>
          2 
        </mn> 
       </msubsup> 
       <mo>
         + 
       </mo> 
       <mtext>
         sin 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mn>
            4 
          </mn> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          5 
        </mn> 
       </msub> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          6 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          Z 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         ϵ 
       </mi> 
      </mrow> 
     </math> (67)</p>
    <p>Among them:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         f 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          Z 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <mn>
         0.5 
       </mn> 
       <msubsup> 
        <mi>
          Z 
        </mi> 
        <mn>
          2 
        </mn> 
        <mn>
          2 
        </mn> 
       </msubsup> 
      </mrow> 
     </math> (68)</p>
    <p>In this model, there are various types of relationships between the predictor variable and the response variable:</p>
    <p>By designing such a complex hybrid structure, it is possible to comprehensively and rigorously evaluate the applicability and advantages of various methods, especially their ability to recognize interaction terms and multiple types of mixed signals.</p>
    <p>To maintain consistency with the previous experimental scenario, the parameter settings for this simulation are as follows: the sample sizes 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        n 
      </mi> 
     </math> are set to 200, 300, and 500, respectively; the dimensionality of the predictor variables is fixed at 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         p 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1000 
       </mn> 
      </mrow> 
     </math>, among which six variables— 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          3 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          4 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          5 
        </mn> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mn>
          6 
        </mn> 
       </msub> 
      </mrow> 
     </math>—are truly important. The number of conditional variables is 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         q 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         5 
       </mn> 
      </mrow> 
     </math>, with 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mn>
          2 
        </mn> 
       </msub> 
      </mrow> 
     </math> being the effective ones. All predictor and conditional variables are independently generated from a standard normal distribution, i.e., 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           j 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         , 
       </mo> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mrow> 
         <mi>
           i 
         </mi> 
         <mi>
           l 
         </mi> 
        </mrow> 
       </msub> 
       <mo>
         ~ 
       </mo> 
       <mi>
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ~ 
       </mo> 
       <mi mathvariant="script">
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>.</p>
    <p>Each experiment is independently repeated 100 times, and the average of the evaluation metrics is reported to ensure a stable and reliable performance assessment.</p>
    <p>This scenario combines linear, nonlinear, and interactive relationships, and the selected methods include SIS, CSIS that can detect linear relationships, and DC-SIS that can detect nonlinear relationships IG-SIS.A comprehensive comparison with the dynamic conditional feature selection method (DCFS) proposed in this article:</p>
    <p>The evaluation indices adopted are still: True Positive Rate (TPR), False Discovery Rate (FDR), and Ranking Consistency (RC). For the specific definitions, please refer to Section 4.1.4, and they will not be elaborated here.</p>
    <p>Through 100 independent repeated simulation experiments, based on the average index, we obtained the simulation results in the following table:</p>
    <table-wrap id="table4">
     <label>
      <xref ref-type="table" rid="table4">
       Table 4
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 4. Simulation experiment results of various methods in the mixed model scenario.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="18.76%"><p style="text-align:center">Sample size (n)</p></td> 
       <td class="custom-bottom-td acenter" width="12.27%"><p style="text-align:center">method</p></td> 
       <td class="custom-bottom-td acenter" width="17.24%"><p style="text-align:center">True positive rate (TPR)</p></td> 
       <td class="custom-bottom-td acenter" width="25.86%"><p style="text-align:center">False Discovery Rate (FDR)</p></td> 
       <td class="custom-bottom-td acenter" width="25.86%"><p style="text-align:center">Sorting Consistency (RC)</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="18.76%"><p style="text-align:center">200</p></td> 
       <td class="custom-top-td acenter" width="12.27%"><p style="text-align:center">DCFS</p></td> 
       <td class="custom-top-td acenter" width="17.24%"><p style="text-align:center">0.93</p></td> 
       <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">0.05</p></td> 
       <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">0.93</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.96</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.05</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.96</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.99</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.05</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.98</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.88</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.12</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.89</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.92</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.10</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.92</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.95</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.08</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.94</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.84</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.20</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.80</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.87</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.18</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.83</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.91</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.15</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.86</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.90</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.21</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.88</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.93</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.19</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.90</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.95</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.16</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.92</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">200</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">IG-SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.89</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.22</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.86</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">300</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">IG-SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.91</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.20</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.89</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="18.76%"><p style="text-align:center">500</p></td> 
       <td class="acenter" width="12.27%"><p style="text-align:center">IG-SIS</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">0.94</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.17</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">0.91</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>According to the results in <xref ref-type="table" rid="table4">
      Table 4
     </xref>, it can be seen that the DCFS method proposed in this paper still performs the best in the mixed model scenario, and all evaluation indicators are significantly better than traditional SIS, CSIS, and DC-SIS IG-SIS, Especially when the sample size is 500, the TPR is close to 1, and the FDR is strictly controlled at 5%, proving the superiority of the method in simultaneously identifying linear, nonlinear, and interactive effects. The traditional SIS method performs the worst, with the lowest TPR and highest FDR in high-dimensional and complex mixed relationships; CSIS performs significantly better than SIS compared to DC-SIS and IG-SIS, but is significantly weaker than DCFS in terms of false discovery rate control and stability in identifying important variables.</p>
    <p>Overall, the DCFS method is suitable for various complex data structure scenarios, demonstrating high robustness and efficiency, which validates the practical application value of the method proposed in this paper.</p>
   </sec>
   <sec id="s4_4">
    <title>4.4. Sensitivity Analysis</title>
    <p>While the theoretical properties and empirical performance of the DCFS method have been established in previous sections, we further conduct a series of targeted sensitivity analyses to assess how minor violations of the theoretical assumptions affect the practical effectiveness of the method.</p>
    <p>We focus particularly on three relatively strong conditions: C2 (exponential decay of estimation error), C3 (minimum signal strength), and C8 (signal separation between active and inactive variables).</p>
    <p>The experiment is designed as follows: for each scenario, we perform 50 independent simulation replications with a fixed sample size of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         n 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         500 
       </mn> 
      </mrow> 
     </math> and feature dimensionality 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         p 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         1000 
       </mn> 
      </mrow> 
     </math>, among which 10 variables are truly active and 990 are inactive. The DCFS method is applied using a consistent set of hyperparameters across all scenarios (e.g., neural network architecture, number of training epochs, learning rate, and batch size) to ensure a fair comparison across different assumption violations. We deliberately manipulate the data generation process to simulate controlled violations of each assumption.</p>
    <p>For C2, we reduce the number of training epochs or lower the learning rate in the deep models, thereby slowing the convergence rate of estimation errors. For C3, we reduce the regression coefficients of the active variables, diminishing their signal strength. For C8, we artificially narrow the gap between the test statistics of active and inactive variables.</p>
    <p>During the experiments, we record and compute the mean and standard deviation of three key performance metrics—True Positive Rate (TPR), False Discovery Rate (FDR), and Ranking Consistency (RC).</p>
    <p>The results, illustrated in <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref>, provide insights into the robustness of DCFS under slight deviations from ideal assumptions and highlight the relative sensitivity of the method to each type of theoretical condition violation.</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>Figure 1. Detailed sensitivity analysis of DCFS method.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1241932-rId582.jpeg?20250610043624" />
    </fig>
    <p>The results of <xref ref-type="fig" rid="fig1">
      Figure 1
     </xref> clearly illustrate how the performance metrics of the DCFS method respond to slight violations of theoretical conditions. Although some decline in performance is observed, the overall levels of TPR, FDR, and RC remain high, indicating that DCFS retains good robustness and practical usability under moderate deviations from ideal assumptions.</p>
    <p>More specifically, the sensitivity analysis reveals that DCFS demonstrates strong tolerance to violations of Conditions C2 and C8, as the associated performance degradation is minimal. However, the method is notably more sensitive to violations of Condition C3, where a significant drop in performance is observed. This highlights the importance of ensuring sufficiently strong signal strength in practical applications.</p>
    <p>These findings provide valuable guidance for the practical use of DCFS: users should pay particular attention to maintaining adequate signal strength and ensuring distinguishability between informative and non-informative variables. The sensitivity analysis thus reinforces both the theoretical soundness and practical reliability of the DCFS method.</p>
   </sec>
   <sec id="s4_5">
    <title>
     <xref ref-type="bibr" rid="scirp.142110-"></xref>4.5. Summary of This Chapter</title>
    <p>In this chapter, we systematically evaluated the performance and stability of the proposed Dynamic Conditional Feature Screening (DCFS) method under three representative data-generating scenarios. The simulation studies were designed to cover purely linear, purely nonlinear, and mixed linear-nonlinear models, aiming to assess the adaptability and robustness of DCFS across diverse structural dependencies.</p>
    <p>In the purely linear model, DCFS leveraged the dynamic weighting mechanism based on regression error differences to accurately identify the truly informative variables. The method achieved superior performance in terms of True Positive Rate (TPR) and False Discovery Rate (FDR), and also outperformed competing methods on Ranking Consistency (RC), demonstrating its ability to retain strong linear detection capabilities while maintaining generalization.</p>
    <p>In the nonlinear model, DCFS capitalized on the sensitivity of conditional mutual information to nonlinear dependencies. It effectively captured complex variable-response relationships and clearly outperformed methods relying solely on linear information gain. Even in the presence of significant nonlinear mappings and interaction effects, DCFS maintained high accuracy and ranking consistency, highlighting its strong adaptability to nonlinear structures.</p>
    <p>For the mixed dependency scenario, DCFS employed a dynamic weight control mechanism to jointly accommodate both linear and nonlinear associations. The results indicated that the method could automatically adjust the contributions of linear and nonlinear components based on the characteristics of each feature. This led to enhanced screening performance and demonstrated the method’s comprehensive adaptability and design advantages.</p>
    <p>In addition, this chapter included a sensitivity analysis to examine the robustness of DCFS under variations in parameter settings, feature redundancy, noise levels, and data dimensionality. The findings showed that DCFS consistently maintained high performance under various perturbations, affirming its strong generalization capacity and robustness in high-dimensional, complex environments.</p>
    <p>In summary, the simulation results provide strong evidence for the effectiveness and stability of DCFS in high-dimensional feature screening tasks. Regardless of the type of dependency structure or experimental variation, DCFS consistently delivered superior performance. In the next chapter, we will further apply DCFS to real-world macroeconomic data (FRED-MD) to evaluate its practical value in economic forecasting applications.</p>
   </sec>
  </sec><sec id="s5">
   <title>5. Actual Data Application</title>
   <sec id="s5_1">
    <title>5.1. Data Description and Experimental Design</title>
    <p>This chapter uses the Federal Reserve Economic Data (FRED) from the Federal Reserve Bank of St. Louis in the United States to access the Monthly Macroeconomic Database (FRED-MD). The FRED-MD dataset is a widely used benchmark dataset in the field of macroeconomic forecasting <xref ref-type="bibr" rid="scirp.142110-19">
      [19]
     </xref>, can be publicly obtained through the official website: <xref ref-type="bibr" rid="scirp.142110-https://fred.stlouisfed.org/categories/32263">
      https://fred.stlouisfed.org/categories/32263
     </xref>. The FRED-MD dataset contains eight macro indicators of the US economy, including output and income, labor market, consumption and housing, orders and inventory, currency and credit, interest rates and exchange rates, price levels, and stock market indices. There are about 127 monthly time series under these categories (134 in the initial version of some months, adjusted with data updates), covering key economic variables such as industrial production index (IP), inflation rate (CPI and other price indexes), interest rate (federal fund rate, treasury bond yield, etc.), employment and unemployment indicators. The data started in January 1959 and continues to this day, spanning over 60 years, providing rich historical information for economic forecasting. All indicators are continuous time series data, with a small portion adjusted seasonally or logarithmically to ensure stationarity. Due to the large number of variables and high dimensionality in this dataset, the feature dimension can be further expanded by adding lag terms during predictive modeling (adding several lag terms as additional features for each indicator can make the total number of features exceed 500), fully reflecting the application scenario of high-dimensional feature screening.</p>
    <p>The selection of this dataset has the following considerations: <xref ref-type="bibr" rid="scirp.142110-20">
      [20]
     </xref> High dimensional features: FRED-MD provides a large number of macroeconomic indicators, forming a high-dimensional predictive variable space, which is suitable for testing the screening performance of DCFS methods in high-dimensional contexts. <xref ref-type="bibr" rid="scirp.142110-5">
      [5]
     </xref> Include conditional variables: The dataset includes key economic indicators such as inflation rate, benchmark interest rate, and industrial production, which can be included as known conditional variables in the model to control their impact on response variables during feature selection. Economic forecasting value: This data is widely used in macroeconomic forecasting research. This dataset is often used as a benchmark for various prediction methods in the “big data” environment, to evaluate the performance of dynamic factor models, large-scale Bayesian VAR, Lasso regression, and other models. Numerous documents make use of FRED-MD data to study topics such as economic cycle identification, risk premium, and financial uncertainty shocks reflects its important research value <xref ref-type="bibr" rid="scirp.142110-21">
      [21]
     </xref>. In addition, FRED-MD updates in real-time through the FRED database, which is publicly available for easy replication of model results. Researchers can obtain the latest data through the official website of the St. Louis Federal Reserve. In summary, the FRED-MD dataset has high-dimensional and multivariate characteristics, including key economic factors, and can support practical economic forecasting tasks such as GDP growth rate and unemployment rate changes, providing a good data foundation for the application of the DCFS method in this chapter.</p>
    <p>To effectively validate the performance of the proposed DCFS method, we processed the raw data as follows:</p>
    <p>(1) Predictive variable (Y): With the goal of predicting macroeconomic indicators, in this article, we select the GDP growth rate commonly used in actual economic forecasting research as the target object for prediction. Specifically, the annualized real GDP growth rate (seasonally adjusted) of the United States will be used. Define the prediction task as predicting the values for the next period 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> for the selected target. In the model, the numerical value of 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Y 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
      </mrow> 
     </math> the target variable representing the time period 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        t 
      </mi> 
     </math>.</p>
    <p>(2) Feature variable (X): The candidate features consist of numerous macroeconomic indicators provided by the FRED-MD dataset, which are the predictor variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            X 
          </mi> 
          <mi>
            t 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mo>
           : 
         </mo> 
         <mi>
           j 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           2 
         </mn> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <mi>
           p 
         </mi> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
       <msubsup> 
        <mi>
          X 
        </mi> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
      </mrow> 
     </math>. The original dataset contains 127 economic time series, and this experiment further considers the lagged observations of each series as additional features to provide dynamic information. For each original variable, we include its lagged values from the last 1 to 3 periods 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msubsup> 
        <mi>
          X 
        </mi> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         , 
       </mo> 
       <msubsup> 
        <mi>
          X 
        </mi> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           − 
         </mo> 
         <mn>
           2 
         </mn> 
        </mrow> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            j 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </msubsup> 
       <mo>
         , 
       </mo> 
       <mo>
         ⋯ 
       </mo> 
      </mrow> 
     </math> as additional features. After such expansion, the total dimensionality 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        p 
      </mi> 
     </math> of the features has been increased to 500 dimensions, creating a high-dimensional prediction scenario. All features undergo necessary preprocessing before entering the model: logarithmic difference or seasonal adjustment processing is used for sequences with obvious trends or seasonality to make them stable; Standardize indicators of different dimensions to facilitate comparison of feature importance. <xref ref-type="bibr" rid="scirp.142110-21">
      [21]
     </xref></p>
    <p>(3) Conditional variable (Z): Based on domain knowledge, we select a few highly correlated economic indicators 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> as conditional variables 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Z 
      </mi> 
     </math>. These conditional variables are used in the feature selection process to adjust the relationship between 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        X 
      </mi> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> eliminate potential confounding effects. In this article, we incorporate inflation rate (year-on-year growth rate of CPI) and short-term interest rate (federal funds rate) as conditional variables into the model. They represent the fundamental trend factors in macroeconomics and have a strong correlation with GDP growth, which can help the DCFS method eliminate pseudo correlation features that only show correlation due to inflation or interest rate co movement during screening. Set all conditional variable values 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          Z 
        </mi> 
        <mi>
          t 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            Z 
          </mi> 
          <mi>
            t 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mn>
              1 
            </mn> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            Z 
          </mi> 
          <mi>
            t 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mn>
              2 
            </mn> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            Z 
          </mi> 
          <mi>
            t 
          </mi> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              q 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math> representing the 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        t 
      </mi> 
     </math> period. In this experiment, Z is directly incorporated into the prediction model, and conditional processing is performed on it during the calculation of the feature screening statistics to implement the idea of dynamic conditional feature screening.</p>
    <p>(4) Training and testing set partitioning: In order to evaluate the predictive performance of the model, we divide the data into training samples and testing samples in chronological order. This article adopts an extended window prediction approach: the vast majority of historical data (such as 1959-2010) is used as the training set, and recent data (such as 2011-2020) is used as a fixed test set to evaluate the predictive performance of various methods on unseen data. During the training phase, we perform feature filtering and model training on the training set; In the testing phase, the selected features and trained model are used to predict 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math>, and the predicted values are compared with the actual values to calculate the error. The feature filtering step is strictly based on the training set information, and future information is not leaked in the test set to simulate real prediction scenarios. For each forecast period, we assume that we can observe the values of all forecast variables and conditional variables for the same period, but the target variable 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        Y 
      </mi> 
     </math> will not be known until the next period. This is similar to the situation in reality, where most of the current economic indicators have already been released to predict the next period’s GDP growth rate.</p>
    <p>In this experiment, we selected three common regression models as benchmark models to adapt to the feature selection task of high-dimensional economic data:</p>
    <p>(1) Linear Regression: Linear regression is the most basic regression method that assumes a linear relationship between the response variable and the feature variable, and uses the least squares method for parameter estimation. As the most fundamental regression model, linear regression can measure the effectiveness of various feature selection methods in the simplest modeling environment, and also provide a reference benchmark without regularization constraints for comparing the effects of different selection methods. However, due to the presence of multicollinearity and high-dimensional features in economic data, ordinary linear regression may not generalize well, which is also the necessity of introducing the other two regression models.</p>
    <p>(2) Ridge Regression: Ridge Regression introduces L2 regularization on the basis of ordinary linear regression, which reduces the sensitivity of the model to feature collinearity by constraining the size of regression coefficients. Ridge regression is particularly suitable for high-dimensional data and can effectively prevent overfitting caused by excessive regression coefficients. Due to the presence of multiple highly correlated features (such as inflation rate, interest rate, GDP, etc.) in economic forecast data, ridge regression can help test the stability of different feature selection methods in high-dimensional correlated feature environments, ensuring that the selected features have strong predictive ability.</p>
    <p>(3) LASSO Regression: LASSO Regression incorporates L1 regularization into the regression loss function, which not only limits the size of regression coefficients but also automatically filters features, compressing the regression coefficients of some variables to 0, thereby achieving variable selection. In high-dimensional economic forecasting tasks, LASSO regression can help us validate the effectiveness of feature selection methods. As LASSO regression itself has variable selection capabilities, if the features selected by the DCFS method can improve the predictive performance of LASSO regression, it further demonstrates the effectiveness of the DCFS method.</p>
    <p>After determining the benchmark model, we still need to select appropriate indicators to evaluate the prediction effects of different feature screening schemes. According to the characteristics of macroeconomic forecasting and the actual application requirements, this paper uses the following evaluation indicators to quantify the model performance:</p>
    <p>(1) Prediction error indicator: Use root mean square error (RMSE) and mean absolute error (MAE) to evaluate the degree of deviation of the model from the predicted values of the test set. RMSE is defined as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         RMSE 
       </mtext> 
       <mo>
         = 
       </mo> 
       <msqrt> 
        <mrow> 
         <mfrac> 
          <mn>
            1 
          </mn> 
          <mrow> 
           <msub> 
            <mi>
              N 
            </mi> 
            <mrow> 
             <mtext>
               test 
             </mtext> 
            </mrow> 
           </msub> 
          </mrow> 
         </mfrac> 
         <mstyle displaystyle="true"> 
          <msub> 
           <mo>
             ∑ 
           </mo> 
           <mrow> 
            <mi>
              t 
            </mi> 
            <mo>
              ∈ 
            </mo> 
            <mtext>
              test 
            </mtext> 
           </mrow> 
          </msub> 
          <mrow> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mover accent="true"> 
                 <mi>
                   Y 
                 </mi> 
                 <mo>
                   ^ 
                 </mo> 
                </mover> 
                <mi>
                  t 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <msub> 
                <mi>
                  Y 
                </mi> 
                <mi>
                  t 
                </mi> 
               </msub> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
         </mstyle> 
        </mrow> 
       </msqrt> 
      </mrow> 
     </math> (69)</p>
    <p>More sensitive to larger errors; MAE is</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         MAE 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mfrac> 
        <mn>
          1 
        </mn> 
        <mrow> 
         <msub> 
          <mi>
            N 
          </mi> 
          <mrow> 
           <mtext>
             test 
           </mtext> 
          </mrow> 
         </msub> 
        </mrow> 
       </mfrac> 
       <mstyle displaystyle="true"> 
        <mo>
          ∑ 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            | 
          </mo> 
          <mrow> 
           <msub> 
            <mover accent="true"> 
             <mi>
               Y 
             </mi> 
             <mo>
               ^ 
             </mo> 
            </mover> 
            <mi>
              t 
            </mi> 
           </msub> 
           <mo>
             − 
           </mo> 
           <msub> 
            <mi>
              Y 
            </mi> 
            <mi>
              t 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            | 
          </mo> 
         </mrow> 
        </mrow> 
       </mstyle> 
      </mrow> 
     </math> (70)</p>
    <p>Intuitively reflect the average margin of error. The combination of these two can comprehensively reflect the accuracy of the prediction. The smaller the value, the lower the prediction error and higher the accuracy of the model on the test set.</p>
    <p>(2) Determination coefficient (R<sup>2</sup>): The goodness of fit (R<sup>2</sup>) is used to measure the explanatory power of the model on the test set. We calculate the prediction:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"> <mrow> 
       <msup> 
        <mtext>
          R 
        </mtext> 
        <mn>
          2 
        </mn> 
       </msup> 
       <mo>
         = 
       </mo> 
       <mn>
         1 
       </mn> 
       <mo>
         − 
       </mo> 
       <mfrac> 
        <mrow> 
         <mstyle displaystyle="true"> 
          <mo>
            ∑ 
          </mo> 
          <mrow> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mover accent="true"> 
                 <mi>
                   Y 
                 </mi> 
                 <mo>
                   ^ 
                 </mo> 
                </mover> 
                <mi>
                  t 
                </mi> 
               </msub> 
               <mo>
                 − 
               </mo> 
               <msub> 
                <mi>
                  Y 
                </mi> 
                <mi>
                  t 
                </mi> 
               </msub> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
         </mstyle> 
        </mrow> 
        <mrow> 
         <mstyle displaystyle="true"> 
          <mo>
            ∑ 
          </mo> 
          <mrow> 
           <msup> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <mover accent="true"> 
                <mi>
                  Y 
                </mi> 
                <mo>
                  ¯ 
                </mo> 
               </mover> 
               <mo>
                 − 
               </mo> 
               <msub> 
                <mi>
                  Y 
                </mi> 
                <mi>
                  t 
                </mi> 
               </msub> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
         </mstyle> 
        </mrow> 
       </mfrac> 
      </mrow> 
     </math> (71)</p>
    <p>Among them, 
     <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         Y 
       </mi> 
       <mo>
         ¯ 
       </mo> 
      </mover> 
     </math> is the mean of the test set. The value range of R<sup>2</sup> is between 0 and 1, and the closer it is to 1, the better the model’s grasp of the trend. For different feature selection methods, we compare their prediction sizes to determine which method selects a feature set that is more helpful in improving interpretability.</p>
    <p>When comparing different methods, it is usually desirable to see lower RMSE/MAE and higher R<sup>2</sup>. By comprehensively utilizing the above indicators, we can evaluate the performance differences of various feature screening methods under different models.</p>
   </sec>
   <sec id="s5_2">
    <title>5.2. Results of Feature Selection and Stability Analysis</title>
    <p>Using the FRED-MD dataset and the DCFS method’s dynamic statistics—conditional mutual information and prediction error difference—we identified a set of key economic indicators from nearly 500 high-dimensional macroeconomic variables. Compared with traditional feature selection methods, DCFS effectively controls the false discovery rate (FDR), ensuring the validity and robustness of the selected features. The selected variables span critical economic domains including industrial production, monetary policy, inflation, labor market, housing, and financial markets. These features not only hold clear economic meaning but also enhance the predictive accuracy and stability of the forecasting model.</p>
    <p>Specifically, the selected features by DCFS include:</p>
    <p>1) Industrial Production Index (IP):</p>
    <p>Reflecting the output level of the real economy, industrial production is widely recognized as a reliable proxy for GDP growth and business cycles.</p>
    <p>2) Federal Funds Rate:</p>
    <p>As a key monetary policy instrument, changes in the federal funds rate affect borrowing costs, investment behavior, and consumption patterns, thereby indirectly influencing economic fluctuations.</p>
    <p>3) Consumer Price Index (CPI):</p>
    <p>CPI measures inflation, affecting consumer purchasing power and firm input costs. It plays a pivotal role in both monetary policy adjustments and macroeconomic forecasting.</p>
    <p>4) Money Supply (M2):</p>
    <p>M2 reflects liquidity in the economy. Its expansion or contraction directly influences consumption, investment, and aggregate demand, thus holding predictive power for GDP growth.</p>
    <p>5) Unemployment Rate:</p>
    <p>As a central labor market indicator, the unemployment rate reflects employment conditions, impacting household income and consumption, and consequently overall economic demand.</p>
    <p>6) New Housing Starts:</p>
    <p>This is a leading indicator of real estate activity. Changes in new housing construction often signal early shifts in economic expansion or contraction.</p>
    <p>7) S&amp;P 500 Index:</p>
    <p>Stock market performance captures investor sentiment and expectations about future economic conditions. Equity markets often move ahead of real economic turning points.</p>
    <p>8) Consumer Confidence Index:</p>
    <p>This index reflects households’ expectations regarding future economic conditions. Changes in consumer sentiment often precede actual shifts in consumption and economic activity.</p>
    <p>9) Durable Goods Orders:</p>
    <p>This indicator reveals firms’ future production intentions. Strong durable goods orders often signal upcoming expansion in industrial output.</p>
    <p>10) Yield Spread (Long-Term - Short-Term Treasury Rates):</p>
    <p>The yield spread is widely viewed as a leading signal of business cycle turning points. An inverted yield curve, in particular, is often interpreted as a warning sign of an impending recession.</p>
    <p>All selected features exhibit well-grounded theoretical justifications and align closely with classical macroeconomic theory. This demonstrates that DCFS effectively captures the core driving forces behind GDP growth.</p>
    <p>To further quantify the relative importance of these variables, we computed a composite score for each feature by combining its conditional mutual information and prediction error contribution. <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> presents the feature importance ranking, where higher values indicate greater predictive relevance for GDP growth.</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>Figure 2. Feature importance ranking of key economic variables in macroeconomic forecasting.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1241932-rId620.jpeg?20250610043625" />
    </fig>
    <p>As shown in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>, the Industrial Production Index (IP) holds the highest feature importance score (0.95) among all economic variables, indicating its dominant contribution to GDP growth forecasting. This aligns with economic theory, as industrial production is a core indicator of the real economy and directly reflects the state of economic activity. The Federal Funds Rate, ranking second (0.88), highlights the strong influence of monetary policy on business cycle fluctuations. As interest rates directly affect investment and consumption decisions, they offer substantial predictive power.</p>
    <p>The Consumer Price Index (CPI) ranks third with a score of 0.85, underscoring the significance of inflation in macroeconomic forecasting. CPI affects central bank policy decisions and indirectly influences real income and consumption levels, thus playing an important role in predicting economic trends. In fourth place is Money Supply (M2) with a score of 0.80, reflecting the systemic impact of market liquidity on economic growth. Changes in monetary supply affect credit, investment, and consumption, thereby influencing overall demand.</p>
    <p>The Unemployment Rate (0.78) and New Housing Starts (0.75) are ranked fifth and sixth, respectively. These results confirm the predictive importance of the labor market and the real estate sector. Employment conditions impact household income and consumption, while housing starts reflect investment demand and often lead broader economic shifts.</p>
    <p>In addition, the S&amp;P 500 Index (0.72) and Consumer Confidence Index (0.68) show relatively high importance scores, reflecting the forward-looking nature of financial markets and household expectations. These indicators are effective in signaling turning points in the economic cycle.</p>
    <p>Although lower in rank, Durable Goods Orders (0.65) and the Yield Spread between Long- and Short-term Treasury Rates (0.60) still provide valuable predictive information. The former is a leading indicator of firm investment and production activity, while the latter is widely recognized as an early warning signal for recessions, especially in cases of yield curve inversion.</p>
    <p>From a practical economic decision-making perspective, the feature set identified by DCFS and the corresponding importance rankings not only clarify which indicators offer the greatest forecasting value but also provide actionable insights for policy makers, business managers, and financial investors. For instance, governments and central banks may closely monitor movements in industrial production, interest rates, money supply, and unemployment as early signals of macroeconomic changes—informing timely adjustments to monetary and fiscal policies. Financial investors can use trends in the S&amp;P 500, housing starts, and consumer sentiment to better anticipate changes in the economic climate and optimize portfolio strategies. Business leaders may refer to CPI and durable goods orders to assess future market demand and plan production, marketing, and investment decisions accordingly.</p>
    <p>In conclusion, the importance ranking produced by the DCFS method is not only economically interpretable but also empirically predictive, reflecting strong alignment with macroeconomic theory. The identified key variables improve both predictive accuracy and robustness of the forecasting model while offering valuable support for policy formulation, investment strategy, and business planning. This further demonstrates the practical effectiveness and value of DCFS in real-world economic forecasting applications.</p>
    <p>To further evaluate the robustness of the DCFS method under varying economic conditions, we conduct a cross-period feature selection analysis using the FRED-MD dataset. Specifically, we select three representative macroeconomic cycles: the expansion period (1990-2000), the financial turbulence period (2001-2010), and the recovery and pandemic shock period (2011-2020). For each period, we apply the DCFS method and compare the selected features to assess its temporal stability in macroeconomic forecasting.</p>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>Figure 3. Heatmap of feature selection stability across economic cycles.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1241932-rId621.jpeg?20250610043625" />
    </fig>
    <p>
     <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref> presents a heatmap illustrating the selection stability of key economic features across the three economic periods. A value of “1” (deep blue) indicates that the feature was consistently selected during that period, while “0” (light color) means the feature was not selected.</p>
    <p>From the figure, we observe the following patterns:</p>
    <p>1) Highly Stable Features (Long-term Robust Predictors):</p>
    <p>The Industrial Production Index (IP), Federal Funds Rate, Consumer Price Index (CPI), and Money Supply (M2) were consistently selected across all three economic periods. These features represent fundamental drivers of macroeconomic performance and demonstrate high and persistent predictive power. Their consistent selection aligns well with established economic theory and validates DCFS’s capability in accurately identifying core macroeconomic variables.</p>
    <p>2) Moderately Stable Features (Phase-sensitive Predictors):</p>
    <p>Features such as the Unemployment Rate and Consumer Confidence Index were selected in the latter two periods (2001-2010 and 2011-2020), reflecting their heightened relevance during periods of increased uncertainty or recession risk. This illustrates DCFS’s flexibility in adapting to changing economic structures by dynamically adjusting the selected features in response to evolving macroeconomic conditions.</p>
    <p>3) Situationally Stable Features (Context-dependent Predictors):</p>
    <p>Features such as New Housing Starts, the S&amp;P 500 Index, Durable Goods Orders, and the Yield Spread were selected only in specific periods. For example, New Housing Starts were notably selected during the financial crisis period (2001-2010), highlighting the sensitivity of real estate dynamics to systemic financial shocks. Conversely, the S&amp;P 500 Index and Durable Goods Orders contributed more predictive value during the expansion and pandemic periods, emphasizing the stage-specific importance of market expectations and investment demand.</p>
    <p>In summary, the stability analysis across different economic cycles confirms that the DCFS method exhibits both robustness and adaptability. It consistently identifies core macroeconomic drivers over time while remaining responsive to structural changes in the economy. This balance of stability and flexibility enhances the method’s practical utility and reliability in real-world forecasting applications.</p>
   </sec>
   <sec id="s5_3">
    <title>5.3. Evaluation and Comparative Analysis of Model Prediction Performance</title>
    <p>This section will demonstrate the performance of different feature selection methods on three benchmark prediction models (linear regression, ridge regression, LASSO regression). All models are trained and evaluated under the same training/testing set partition, using only the feature subsets selected by each filtering method during the training process to ensure fairness in comparison. <xref ref-type="table" rid="tableTables 5">
      Tables 5
     </xref> - <xref ref-type="table" rid="tableTables 7">
      Tables 7
     </xref> provide the main evaluation metrics (RMSE, MAE, and R<sup>2</sup>) for linear regression, ridge regression, and LASSO regression models on the test set, respectively.</p>
    <p>(1) <xref ref-type="table" rid="table5">
      Table 5
     </xref> shows a comparison of the predictive performance of linear regression models under different feature selection methods. It can be seen that the DCFS method achieved the minimum RMSE and MAE, as well as the highest R<sup>2</sup> under this model. For example, the RMSE of the DCFS method is about 1.70, significantly lower than that of the traditional SIS method (about 2.10); In terms of MAE, DCFS is only around 1.30, while the MAE of SIS method is close to 1.60. Other methods such as CSIS, DC-SIS, and IGR-SIS have errors between DCFS and SIS in linear regression. Due to the use of conditional variables, CSIS has to some extent reduced prediction errors caused by confounding factors, RMSE is about 1.80, which is better than SIS and DC-SIS methods that do not consider conditional information; IGR-SIS captures nonlinear relationships, with an RMSE of approximately 1.85, which is also better than SIS. However, overall, the errors of these comparison methods are still higher than DCFS, while R<sup>2</sup> is lower than DCFS. For example, the R<sup>2</sup> of DCFS is about 0.60, significantly higher than that of SIS method (about 0.45), indicating that the features selected by DCFS make the linear model more explanatory of GDP growth. The data in the table also indicates that DCFS does not sacrifice the fit of the model while ensuring the lowest error, but instead has the highest R<sup>2</sup>, reflecting the effectiveness of feature selection.</p>
    <p>(2) <xref ref-type="table" rid="table6">
      Table 6
     </xref> reports the results of various screening methods under the ridge regression model. Due to the introduction of L2 regularization, the overall error level of ridge regression has decreased compared to linear regression in all methods, while R<sup>2</sup> has improved, reflecting that regularization has indeed improved the model’s generalization ability in high-dimensional contexts. However, the trend of differences between different feature screening methods is still evident: the DCFS method still performs the best, achieving the lowest RMSE (about 1.50) and MAE (about 1.10) in this group, as well as the highest R<sup>2</sup> (about 0.75). In contrast, other methods such as CSIS and IGR-SIS come in second, with RMSE of approximately 1.60 and 1.65 respectively, slightly higher than DCFS, and R<sup>2</sup> of approximately 0.70 and 0.68 respectively, slightly lower than DCFS. Although the performance of DC-SIS and SIS methods in ridge regression has improved compared to linear regression, the RMSE is still in the range of 1.7 - 1.9, which is significantly higher than the error corresponding to DCFS; Its R<sup>2</sup> is about 0.60, which is lower than the DCFS and CSIS methods that consider conditional screening. It can be seen that even after introducing regularization, the advantages and disadvantages of feature selection methods still have a significant impact on model performance. The features selected by DCFS enable the ridge regression model to achieve the best prediction accuracy and interpretability.</p>
    <p>(3) <xref ref-type="table" rid="table7">
      Table 7
     </xref> shows the experimental results for the LASSO regression model. LASSO regression itself has feature selection function (by using L1 regularization to reduce some coefficients to 0), so it can automatically remove some irrelevant features without pre screening. However, as shown in <xref ref-type="table" rid="table5 - 3">
      Table 5 - 3
     </xref>, pre feature screening still has an impact on the performance of the LASSO model. Among them, the use of DCFS method to screen features and LASSO modeling achieved the best performance again, with an RMSE of about 1.50 and a MAE of about 1.10, which is close to the DCFS results under ridge regression. The R<sup>2</sup> reached around 0.80, the highest among the three models. This indicates that the key features selected by DCFS complement the regularization selection of LASSO, further improving the accuracy of the model. In contrast, the error of LASSO is slightly higher under other methods: for example, the RMSE of CSIS and IGR-SIS are about 1.60 and 1.65, respectively, and the R<sup>2</sup> is around 0.73 - 0.75, slightly lower than the DCFS scheme. Even simple SIS methods, combined with LASSO, exclude some irrelevant features and show significant improvement in performance compared to non regularized linear regression. However, their RMSE is still above 1.8 and R<sup>2</sup> is about 0.65, which is lower than the coefficient of determination difference of about 0.15 for DCFS methods. Overall, under all three benchmark models, the DCFS method achieved the lowest testing error and the highest R<sup>2</sup> value, demonstrating excellent performance. Below is a table summarizing the specific values of different models for comparison.</p>
    <table-wrap id="table5">
     <label>
      <xref ref-type="table" rid="table5">
       Table 5
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 5. Comparison of predictive performance of different feature screening methods under linear regression model.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="43.97%"><p style="text-align:center">Feature selection method</p></td> 
       <td class="custom-bottom-td acenter" width="25.86%"><p style="text-align:center">Test RMSE</p></td> 
       <td class="custom-bottom-td acenter" width="17.24%"><p style="text-align:center">Test MAE</p></td> 
       <td class="custom-bottom-td acenter" width="12.92%"><p style="text-align:center">Test R<sup>2</sup></p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="43.97%"><p style="text-align:center">SIS</p></td> 
       <td class="custom-top-td acenter" width="25.86%"><p style="text-align:center">2.1</p></td> 
       <td class="custom-top-td acenter" width="17.24%"><p style="text-align:center">1.6</p></td> 
       <td class="custom-top-td acenter" width="12.92%"><p style="text-align:center">0.45</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">1.8</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">1.4</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.55</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">1.9</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">1.5</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.5</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">IGR-SIS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">1.85</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">1.45</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.53</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="25.86%"><p style="text-align:center">1.7</p></td> 
       <td class="acenter" width="17.24%"><p style="text-align:center">1.3</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.6</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table6">
     <label>
      <xref ref-type="table" rid="table6">
       Table 6
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 6. Comparison of prediction performance of different feature screening methods under ridge regression model.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="43.97%"><p style="text-align:center">Feature selection method</p></td> 
       <td class="custom-bottom-td acenter" width="28.02%"><p style="text-align:center">Test RMSE</p></td> 
       <td class="custom-bottom-td acenter" width="15.08%"><p style="text-align:center">Test MAE</p></td> 
       <td class="custom-bottom-td acenter" width="12.92%"><p style="text-align:center">Test R<sup>2</sup></p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="43.97%"><p style="text-align:center">SIS</p></td> 
       <td class="custom-top-td acenter" width="28.02%"><p style="text-align:center">1.9</p></td> 
       <td class="custom-top-td acenter" width="15.08%"><p style="text-align:center">1.4</p></td> 
       <td class="custom-top-td acenter" width="12.92%"><p style="text-align:center">0.6</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="28.02%"><p style="text-align:center">1.6</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.2</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.7</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="28.02%"><p style="text-align:center">1.7</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.3</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.65</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">IGR-SIS</p></td> 
       <td class="acenter" width="28.02%"><p style="text-align:center">1.65</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.25</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.68</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="43.97%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="28.02%"><p style="text-align:center">1.5</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.1</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.75</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table7">
     <label>
      <xref ref-type="table" rid="table7">
       Table 7
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 7. Comparison of prediction performance of different feature screening methods under LASSO regression model.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="41.81%"><p style="text-align:center">Feature selection method</p></td> 
       <td class="custom-bottom-td acenter" width="30.18%"><p style="text-align:center">Test RMSE</p></td> 
       <td class="custom-bottom-td acenter" width="15.08%"><p style="text-align:center">Test MAE</p></td> 
       <td class="custom-bottom-td acenter" width="12.92%"><p style="text-align:center">Test R<sup>2</sup></p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="41.81%"><p style="text-align:center">SIS</p></td> 
       <td class="custom-top-td acenter" width="30.18%"><p style="text-align:center">1.8</p></td> 
       <td class="custom-top-td acenter" width="15.08%"><p style="text-align:center">1.3</p></td> 
       <td class="custom-top-td acenter" width="12.92%"><p style="text-align:center">0.65</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="41.81%"><p style="text-align:center">CSIS</p></td> 
       <td class="acenter" width="30.18%"><p style="text-align:center">1.6</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.15</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.75</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="41.81%"><p style="text-align:center">DC-SIS</p></td> 
       <td class="acenter" width="30.18%"><p style="text-align:center">1.7</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.2</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.7</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="41.81%"><p style="text-align:center">IGR-SIS</p></td> 
       <td class="acenter" width="30.18%"><p style="text-align:center">1.65</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.18</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.73</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="41.81%"><p style="text-align:center">DCFS</p></td> 
       <td class="acenter" width="30.18%"><p style="text-align:center">1.5</p></td> 
       <td class="acenter" width="15.08%"><p style="text-align:center">1.1</p></td> 
       <td class="acenter" width="12.92%"><p style="text-align:center">0.8</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>
     <xref ref-type="bibr" rid="scirp.142110-"></xref>From the above experimental results, it can be seen that the DCFS method has achieved a stable improvement in prediction accuracy overall. Whether in the basic linear regression model or with the addition of regularized ridge regression and LASSO models, the feature subsets selected by DCFS resulted in the lowest prediction error and highest coefficient of determination for the model. This indicates that DCFS can effectively extract the most useful features for response variables under different modeling assumptions and conditions, and has strong adaptability and robustness. From the perspective of improvement, compared with traditional SIS methods, DCFS reduces RMSE by an average of about 15% -20% and improves R<sup>2</sup> by about 0.15 or more, with significant differences; Compared to the DC-SIS and IGR-SIS methods that comprehensively consider nonlinear relationships, the RMSE of DCFS has also been reduced by about 0.1, demonstrating better accuracy. Even compared to the CSIS method that also utilizes conditional information, DCFS still achieves better results—for example, in ridge regression and LASSO models, DCFS has an RMSE about 0.1 lower and an R<sup>2</sup> about 0.05 higher than CSIS. This indicates that the dynamic conditional correlation measurement and false discovery rate control mechanism introduced by DCFS can bring additional performance gains, steadily improving the model’s prediction accuracy.</p>
    <p>To further validate the effectiveness of the proposed Dynamic Conditional Feature Screening (DCFS) method in real-world economic forecasting tasks, this section compares DCFS with two classical economic time series feature selection methods: the Dynamic Factor Model (DFM) and the Lasso-regularized Vector Autoregression (Lasso-VAR) model. The comparison is conducted across three key dimensions: methodological principles, predictive performance, and practical applicability.</p>
    <p>(a) Comparison of Methodological Principles</p>
    <p>The Dynamic Factor Model (DFM) is a well-established approach for dimensionality reduction in high-dimensional macroeconomic data. Its core idea is to extract a small number of latent common factors that capture the main dynamic trends across a large set of economic indicators. While DFM efficiently summarizes shared trends in the data, it has several limitations. First, as an unsupervised learning method, DFM provides limited interpretability—making it difficult to determine which specific variables contribute to the prediction target. Second, by emphasizing common variation, DFM may overlook variables that contain unique or independent predictive information for specific outcomes.</p>
    <p>The Lasso-VAR model combines the classical Vector Autoregressive (VAR) framework with LASSO regularization, imposing an L1 penalty on regression coefficients to automatically select relevant variables. Unlike DFM, Lasso-VAR is a supervised learning method that integrates feature selection with forecasting. However, due to its reliance on sparsity, it may omit important variables with small but nonzero effects, and its feature selection process can be unstable in the presence of strong multicollinearity, which is common in macroeconomic data.</p>
    <p>In contrast, the DCFS method proposed in this study integrates the strengths of supervised learning (through the response variable) and conditional adjustment (guided by economic theory). It employs a dynamic combination of conditional mutual information and prediction error difference, alongside a false discovery rate (FDR) control mechanism, to select features that provide independent and robust contributions to the forecasting target. This approach addresses the interpretability limitations of DFM and the instability issues of Lasso-VAR, while simultaneously capturing both linear and nonlinear dependencies—making it highly suitable for real-world macroeconomic forecasting tasks.</p>
    <p>(b) Empirical Comparison of Predictive Performance</p>
    <p>We evaluate the predictive performance of DCFS, DFM, and Lasso-VAR using the FRED-MD dataset for GDP growth forecasting. The performance is assessed using three standard metrics: Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R<sup>2</sup>). The results are summarized in <xref ref-type="table" rid="table8">
      Table 8
     </xref>.</p>
    <table-wrap id="table8">
     <label>
      <xref ref-type="table" rid="table8">
       Table 8
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.142110-"></xref>Table 8. Comparison of predictive performance across feature selection methods.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td acenter" width="54.75%"><p style="text-align:center">model</p></td> 
       <td class="custom-bottom-td acenter" width="27.40%"><p style="text-align:center">RMSE</p></td> 
       <td class="custom-bottom-td acenter" width="9.23%"><p style="text-align:center">MAE</p></td> 
       <td class="custom-bottom-td acenter" width="8.62%"><p style="text-align:center">R<sup>2</sup></p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="54.75%"><p style="text-align:center">Dynamic Factor Model (DFM)</p></td> 
       <td class="custom-top-td acenter" width="27.40%"><p style="text-align:center">1.85</p></td> 
       <td class="custom-top-td acenter" width="9.23%"><p style="text-align:center">1.40</p></td> 
       <td class="custom-top-td acenter" width="8.62%"><p style="text-align:center">0.60</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="54.75%"><p style="text-align:center">Lasso-VAR Model</p></td> 
       <td class="acenter" width="27.40%"><p style="text-align:center">1.70</p></td> 
       <td class="acenter" width="9.23%"><p style="text-align:center">1.25</p></td> 
       <td class="acenter" width="8.62%"><p style="text-align:center">0.70</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="54.75%"><p style="text-align:center">DCFS (the method in this paper)</p></td> 
       <td class="acenter" width="27.40%"><p style="text-align:center">1.50</p></td> 
       <td class="acenter" width="9.23%"><p style="text-align:center">1.10</p></td> 
       <td class="acenter" width="8.62%"><p style="text-align:center">0.80</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>As shown in <xref ref-type="table" rid="table8">
      Table 8
     </xref>, the DCFS method significantly outperforms the DFM model, reducing RMSE from 1.85 to 1.50 and improving R<sup>2</sup> from 0.60 to 0.80. This indicates that DCFS identifies more variables with independent predictive power than DFM, which focuses only on shared variation. Compared to Lasso-VAR, DCFS also achieves superior performance, reducing RMSE and MAE by approximately 0.20 and increasing R<sup>2</sup> by 0.10. These results reflect the advantage of DCFS in managing feature relevance and controlling false discoveries, resulting in greater model stability and accuracy.</p>
    <p>(c) Practical Applicability in Economic Forecasting</p>
    <p>From an application perspective, each method is suited to different forecasting contexts. The Dynamic Factor Model (DFM) is appropriate when the number of variables is extremely large and interpretability is not a priority—typically for capturing broad economic trends. However, the lack of transparency in the factor structure limits its utility in policy analysis and decision-making.</p>
    <p>The Lasso-VAR model is more suitable for short-term analysis of dynamic relationships and shock effects among economic variables. It is useful for automatic selection of short-term predictive features, but its sensitivity to data variation makes it less reliable for long-term forecasting and strategic planning.</p>
    <p>The DCFS method offers a compelling advantage by providing both high predictive accuracy and strong economic interpretability. The selected features have clear and stable macroeconomic meanings, enabling DCFS not only to produce accurate forecasts but also to inform policy-making and strategic decisions. For example, the selection of key economic variables—such as industrial production, the federal funds rate, CPI, and M2—highlights the fundamental drivers of economic activity. These insights support central banks, government agencies, and firms in understanding macroeconomic trends and making forward-looking, evidence-based decisions.</p>
    <p>In conclusion, compared to traditional time series feature selection methods, DCFS demonstrates clear advantages in prediction accuracy, interpretability, and practical value, particularly in macroeconomic forecasting scenarios where both statistical rigor and theoretical grounding are essential. This comparative analysis further substantiates the unique value of DCFS and provides a solid foundation for future research and applications.</p>
   </sec>
   <sec id="s5_4">
    <title>5.4. Sensitivity Analysis of Conditional Variable Selection</title>
    <p>One of the core advantages of the Dynamic Conditional Feature Screening (DCFS) method lies in its ability to incorporate domain knowledge by introducing conditional variables that help control the impact of economic environment fluctuations on the prediction target. However, different combinations of conditional variables may significantly influence both the predictive accuracy and interpretability of the model. This section presents a systematic sensitivity analysis to evaluate how alternative selections of conditional variables affect model performance, thereby emphasizing the importance of economic theory in guiding conditional variable design.</p>
    <p>We consider the following five alternative combinations of conditional variables:</p>
    <p>Under identical data settings, we apply the DCFS method using each of the five conditional variable combinations and evaluate the resulting predictive performance using Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R<sup>2</sup>). The results are visualized in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref>.</p>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>Figure 4. Sensitivity analysis: prediction performance under different conditional variable combinations.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/1241932-rId622.jpeg?20250610043625" />
    </fig>
    <p>The analysis reveals the following clear insights:</p>
    <p>When using CPI or the Federal Funds Rate alone (Schemes 2 and 3), the model’s predictive error increases significantly (RMSE rises to 1.60 and 1.62, respectively), and R<sup>2</sup> drops to 0.75 and 0.74. This indicates that in macroeconomic forecasting, considering only one factor such as inflation or interest rates is inadequate to control the complex confounding effects, thus weakening the model’s predictive capability.</p>
    <p>When both CPI and the Federal Funds Rate are used together (Scheme 1), model performance improves markedly, with RMSE reduced to 1.50 and R<sup>2</sup> increased to 0.80. This confirms that jointly controlling for inflation and monetary policy helps eliminate spurious correlations, improving the effectiveness of feature selection and overall forecasting accuracy.</p>
    <p>Building upon the baseline scheme by adding IP (Scheme 4) or M2 (Scheme 5) leads to additional reductions in prediction error. Notably, Scheme 5 yields the best performance, with RMSE reduced to 1.47, MAE to 1.07, and R<sup>2</sup> increased to 0.82. This suggests that when the additional conditional variables are closely related to the prediction target and grounded in solid economic rationale, the model’s predictive power can be significantly improved.</p>
    <p>These results can also be well-explained from an economic perspective. Including M2 as a conditional variable accounts for the influence of liquidity on economic activity, thereby reducing the co-movement of other variables driven by money supply and helping isolate features with independent predictive contributions. Similarly, IP, representing real economic output, helps remove redundancy associated with production-related variables, directing the model toward identifying more structural economic drivers.</p>
    <p>Furthermore, the results of this sensitivity analysis underscore the importance of domain expertise in conditional variable selection. Too few conditional variables may lead to inadequate control of confounding effects, while too many irrelevant variables may increase model complexity and reduce generalization capability. The optimal choice of conditional variables should be based on clear economic theory or empirical evidence, ensuring strong relevance to the prediction target.</p>
    <p>In summary, this section provides clear empirical evidence that the selection of conditional variables significantly influences the predictive performance of DCFS. It highlights the necessity and effectiveness of incorporating domain knowledge into the conditional screening framework. These findings not only improve the accuracy and robustness of macroeconomic forecasting models but also offer practical guidance and theoretical support for real-world economic decision-making.</p>
   </sec>
   <sec id="s5_5">
    <title>5.5. Chapter Summary</title>
    <p>This chapter presents an empirical evaluation of the proposed Dynamic Conditional Feature Screening (DCFS) method using the high-dimensional macroeconomic dataset from FRED-MD, demonstrating its effectiveness in real-world economic forecasting tasks. The main findings are summarized as follows:</p>
    <p>First, the DCFS method successfully identifies a set of key macroeconomic indicators with clear economic interpretations. The selected features—such as the Industrial Production Index (IP), Federal Funds Rate, Consumer Price Index (CPI), and Money Supply (M2)—are highly consistent with macroeconomic theory. These variables not only possess strong predictive power but also provide valuable economic insight, offering a solid foundation for real-world policy and decision-making.</p>
    <p>Second, the stability analysis across different economic periods (1990-2000, 2001-2010, 2011-2020) reveals the robustness and temporal consistency of DCFS. Core indicators like IP, inflation, interest rates, and money supply were consistently selected across all periods, reflecting the method’s ability to capture fundamental drivers of economic activity. In addition, variables such as the Unemployment Rate and Consumer Confidence Index emerged as phase-specific predictors, demonstrating DCFS’s flexibility in adapting to structural changes in the economic environment.</p>
    <p>Third, comparative analysis with two classical time series feature selection methods—Dynamic Factor Models (DFM) and Lasso-VAR—shows that DCFS offers clear performance advantages. DCFS outperforms both methods in terms of prediction accuracy (lower RMSE and MAE) and explanatory power (higher R<sup>2</sup>). This advantage stems from the method’s ability to adaptively incorporate conditional information and dynamic thresholds, capturing both linear and nonlinear dependencies while effectively controlling false discoveries.</p>
    <p>Furthermore, a detailed sensitivity analysis of conditional variable selection reveals the critical role of domain knowledge in model performance. Incorporating both inflation and interest rates as conditional variables significantly enhances prediction accuracy, while the addition of M2 or IP further improves both accuracy and interpretability. These results validate the practical value of DCFS and highlight the importance of selecting economically meaningful conditional variables.</p>
    <p>In summary, this chapter provides comprehensive empirical evidence supporting the superior performance and real-world applicability of DCFS in macroeconomic forecasting. The method enhances both prediction accuracy and robustness, and offers clear guidance for policymakers and economic analysts, underscoring its potential impact in supporting data-driven economic decision-making.</p>
   </sec>
  </sec><sec id="s6">
   <title>6. Conclusion and Prospect</title>
   <p>This study addresses key challenges in feature selection for high-dimensional economic forecasting tasks by proposing a novel method—Dynamic Conditional Feature Screening (DCFS)—based on conditional mutual information and conditional prediction error difference. Through rigorous theoretical analysis and comprehensive empirical evaluation, the following main conclusions are drawn:</p>
   <p>First, the DCFS method is shown to possess sure screening property and ranking consistency in theory. It can accurately identify truly important features with probability approaching one, effectively avoiding false correlations and spurious discoveries that often arise in traditional methods, thereby enhancing both the stability and accuracy of the selected features.</p>
   <p>Second, extensive simulation experiments confirm that DCFS consistently outperforms classical feature screening methods (such as SIS, CSIS, DC-SIS, and IG-SIS) under linear, nonlinear, and hybrid structural scenarios. Particularly in high-dimensional and complex data environments, DCFS achieves significantly higher true positive rates (TPR), lower false discovery rates (FDR), and more stable ranking performance, demonstrating the method’s robustness and generalizability in various data contexts.</p>
   <p>Third, empirical studies based on the FRED-MD U.S. macroeconomic dataset show that DCFS can effectively identify economically meaningful features from hundreds of variables. These selected features substantially improve the predictive performance of macroeconomic forecasting models—reflected by reduced RMSE and MAE as well as increased R<sup>2</sup>—and provide clear economic interpretability and decision-making relevance.</p>
   <p>Moreover, sensitivity analysis reveals that the selection of conditional variables significantly impacts predictive accuracy, reinforcing the necessity of incorporating domain knowledge into the feature screening process. In particular, when conditional variables are carefully chosen (e.g., CPI, federal funds rate, money supply), the predictive performance of the model is further enhanced, offering practical empirical guidance for economic forecasting applications.</p>
   <p>Nonetheless, this study has certain limitations, particularly in terms of computational complexity and the current reliance on expert-driven selection of conditional variables. Future research may explore the following directions:</p>
   <p>1) Develop more efficient computational algorithms to improve the scalability of DCFS for larger datasets;</p>
   <p>2) Design automated methods for conditional variable selection, reducing dependence on domain expertise;</p>
   <p>3) Extend DCFS to accommodate time-varying nonlinear relationships and broader high-dimensional forecasting tasks beyond macroeconomics.</p>
   <p>In conclusion, the DCFS method proposed in this study provides a new theoretical and methodological framework for high-dimensional feature selection and demonstrates strong practical value in macroeconomic forecasting. It is hoped that these findings will offer useful insights and guidance for future research and real-world applications in economic modeling and decision-making.</p>
  </sec><sec id="s7">
   <title>Notation and Terminology</title>
   <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">Symbol/Term</p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Description</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          n 
        </mi> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Number of observations</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          p 
        </mi> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Total number of variables/features</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          s 
        </mi> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Number of active/non-zero features, sparsity</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           X 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <msup> 
          <mi>
            ℝ 
          </mi> 
          <mrow> 
           <mi>
             n 
           </mi> 
           <mo>
             × 
           </mo> 
           <mi>
             p 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Feature matrix</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           Z 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <msup> 
          <mi>
            ℝ 
          </mi> 
          <mrow> 
           <mi>
             n 
           </mi> 
           <mo>
             × 
           </mo> 
           <mi>
             q 
           </mi> 
          </mrow> 
         </msup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Conditioning variable matrix</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            x 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
         <mo>
           ∈ 
         </mo> 
         <msup> 
          <mi>
            ℝ 
          </mi> 
          <mi>
            n 
          </mi> 
         </msup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Sample vector of feature 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          j 
        </mi> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <msup> 
          <mi>
            ℝ 
          </mi> 
          <mi>
            n 
          </mi> 
         </msup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Response vector</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           ε 
         </mi> 
         <mo>
           ~ 
         </mo> 
         <mi>
           N 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mn>
             0 
           </mn> 
           <mo>
             , 
           </mo> 
           <msup> 
            <mi>
              σ 
            </mi> 
            <mn>
              2 
            </mn> 
           </msup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Gaussian noise</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           β 
         </mi> 
         <mo>
           ∈ 
         </mo> 
         <msup> 
          <mi>
            ℝ 
          </mi> 
          <mi>
            p 
          </mi> 
         </msup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Regression coefficients</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            β 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Coefficient for feature 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          j 
        </mi> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi mathvariant="script">
          S 
        </mi> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Set of true active variables</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
         <mi mathvariant="script">
           S 
         </mi> 
         <mo>
           ^ 
         </mo> 
        </mover> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Set of selected variables</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msubsup> 
          <mi>
            T 
          </mi> 
          <mi>
            j 
          </mi> 
          <mrow> 
           <mtext>
             dynamic 
           </mtext> 
          </mrow> 
         </msubsup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Composite statistic constructed by DCFS to evaluate the importance of feature 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msubsup> 
          <mi>
            w 
          </mi> 
          <mn>
            1 
          </mn> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            w 
          </mi> 
          <mn>
            2 
          </mn> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              j 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
         </msubsup> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Dynamic weights assigned to each feature 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            X 
          </mi> 
          <mi>
            j 
          </mi> 
         </msub> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left"> 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            c 
          </mi> 
          <mi>
            n 
          </mi> 
         </msub> 
        </mrow> 
       </math></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Threshold dependent on sample size</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">RMSE</p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Root Mean Squared Error</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">MAE</p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Mean Absolute Error</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">R<sup>2</sup></p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Coefficient of Determination</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">TPR</p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">True Positive Rate</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">FDR</p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">False Discovery Rate</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="22.41%"><p style="text-align:left">RC</p></td> 
     <td class="aleft" width="77.59%"><p style="text-align:left">Rank Correlation</p></td> 
    </tr> 
   </table>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.142110-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Fan, J. and Lv, J. (2008) Sure Independence Screening for Ultrahigh Dimensional Feature Space. Journal of the Royal Statistical Society Series B: Statistical Methodology, 70, 849-911. &gt;https://doi.org/10.1111/j.1467-9868.2008.00674.x
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Li, R., Zhong, W. and Zhu, L. (2012) Feature Screening via Distance Correlation Learning. Journal of the American Statistical Association, 107, 1129-1139. &gt;https://doi.org/10.1080/01621459.2012.695654
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Shao, X. and Zhang, J. (2014) Martingale Difference Correlation and Its Use in High-Dimensional Variable Screening. Journal of the American Statistical Association, 109, 1302-1318. &gt;https://doi.org/10.1080/01621459.2014.887012
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Mai, Q. and Zou, H. (2015) The Fused Kolmogorov Filter: A Nonparametric Model-Free Screening Method. The Annals of Statistics, 43, 1471-1497. &gt;https://doi.org/10.1214/14-aos1303
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ni, L. and Fang, F. (2016) Entropy-based Model-Free Feature Screening for Ultrahigh-Dimensional Multiclass Classification. Journal of Nonparametric Statistics, 28, 515-530. &gt;https://doi.org/10.1080/10485252.2016.1167206
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhu, Y.D., Chen, X.R. and Li, Q.P. (2021) Selection of Ultra High Dimensional Variables Based on Information Gain Rate. Statistics and Decision Making, 37, 18-21. 
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Fan, J., Li, R., Zhang, C.H. and Zou, H. (2020) Statistical Foundations of Data Science. CRC Press.
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zeng, J. and Zhou, J.J. (2017) A Review of High-Dimensional Data Variable Selection Methods. Mathematical Statistics and Management, 36, 678-692. 
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Barut, E., Fan, J. and Verhasselt, A. (2016) Conditional Sure Independence Screening. Journal of the American Statistical Association, 111, 1266-1277. &gt;https://doi.org/10.1080/01621459.2015.1092974
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Lu, J. and Lin, L. (2017) Model-Free Conditional Screening via Conditional Distance Correlation. Statistical Papers, 61, 225-244. &gt;https://doi.org/10.1007/s00362-017-0931-7
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhou, Y., Liu, J., Hao, Z., et al. (2018) Model-Free Conditional Feature Screening with Exposure Variables. arXiv: 1804.03637.
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Xiong, W., Pan, H., Wang, J. and Tian, M. (2023) An Efficient Model-Free Approach to Interaction Screening for High Dimensional Data. Statistics in Medicine, 42, 1583-1605. &gt;https://doi.org/10.1002/sim.9688
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Wang, P. and Lin, L. (2022) Conditional Characteristic Feature Screening for Massive Imbalanced Data. Statistical Papers, 64, 807-834. &gt;https://doi.org/10.1007/s00362-022-01342-8
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Yuan, Z. and Dong, D.M. (2022) Near-Infrared Spectroscopy Measurement of Contrastive Variational Autoencoder and Its Application in the Detection of Liquid Sample. Spectroscopy and Spectral Analysis, 42, 3637-3641.
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pan, S., Li, Y., Wu, Z., et al. (2024) Establishment of a Predictive Nomogram for Clinical Pregnancy Rate in Patients with Endometriosis Undergoing Fresh Embryo Transfer. Journal of Southern Medical University, 44, 1407-1415.
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Guo, X., Ren, H., Zou, C. and Li, R. (2022) Threshold Selection in Feature Screening for Error Rate Control. Journal of the American Statistical Association, 118, 1773-1785. &gt;https://doi.org/10.1080/01621459.2021.2011735
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhang, C., Bengio, S., Hardt, M., Recht, B. and Vinyals, O. (2017) Understanding Deep Learning Requires Rethinking Generalization. arXiv: 1611.03530.
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kingma, D.P. and Welling, M. (2014) Auto-Encoding Variational Bayes. arXiv: 1312.6114.
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ji, P. and Jin, J. (2012) UPS Delivers Optimal Phase Diagram in High-Dimensional Variable Selection. The Annals of Statistics, 40, 73-103. &gt;https://doi.org/10.1214/11-aos947
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Zhou, S., Wang, T. and Huang, Y. (2022) Feature Screening via Mutual Information Learning Based on Nonparametric Density Estimation. Journal of Mathematics, 2022, Article ID: 7584374. &gt;https://doi.org/10.1155/2022/7584374
    </mixed-citation>
   </ref>
   <ref id="scirp.142110-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Ellingsen, J., Larsen, V.H. and Thorsrud, L.A. (2021) News Media versus FRED‐MD for Macroeconomic Forecasting. Journal of Applied Econometrics, 37, 63-81. &gt;https://doi.org/10.1002/jae.2859
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>