<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article">
 <front>
  <journal-meta>
   <journal-id journal-id-type="publisher-id">
    jilsa
   </journal-id>
   <journal-title-group>
    <journal-title>
     Journal of Intelligent Learning Systems and Applications
    </journal-title>
   </journal-title-group>
   <issn pub-type="epub">
    2150-8402
   </issn>
   <issn publication-format="print">
    2150-8410
   </issn>
   <publisher>
    <publisher-name>
     Scientific Research Publishing
    </publisher-name>
   </publisher>
  </journal-meta>
  <article-meta>
   <article-id pub-id-type="doi">
    10.4236/jilsa.2025.174015
   </article-id>
   <article-id pub-id-type="publisher-id">
    jilsa-146372
   </article-id>
   <article-categories>
    <subj-group subj-group-type="heading">
     <subject>
      Articles
     </subject>
    </subj-group>
    <subj-group subj-group-type="Discipline-v2">
     <subject>
      Computer Science 
     </subject>
     <subject>
       Communications
     </subject>
    </subj-group>
   </article-categories>
   <title-group>
    CASCADE-Net: Causality-Aware Spatio-Temporal Dynamics Encoding for Prognostic Prediction in Mild Cognitive Impairment
   </title-group>
   <contrib-group>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Samuel
      </surname>
      <given-names>
       Ocen
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref> 
     <xref ref-type="aff" rid="aff2"> 
      <sup>2</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Lawrence
      </surname>
      <given-names>
       Muchemi
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
    <contrib contrib-type="author" xlink:type="simple">
     <name name-style="western">
      <surname>
       Michaelina Almaz
      </surname>
      <given-names>
       Yohannis
      </given-names>
     </name> 
     <xref ref-type="aff" rid="aff1"> 
      <sup>1</sup>
     </xref>
    </contrib>
   </contrib-group> 
   <aff id="aff1">
    <addr-line>
     aDepartment of Computing and Informatics, University of Nairobi, Nairobi, Kenya
    </addr-line> 
   </aff> 
   <aff id="aff2">
    <addr-line>
     aDepartment of Computer Science, Mountains of the Moon University, Fort Portal, Uganda
    </addr-line> 
   </aff> 
   <pub-date pub-type="epub">
    <day>
     14
    </day> 
    <month>
     10
    </month>
    <year>
     2025
    </year>
   </pub-date> 
   <volume>
    17
   </volume> 
   <issue>
    04
   </issue>
   <fpage>
    237
   </fpage>
   <lpage>
    256
   </lpage>
   <history>
    <date date-type="received">
     <day>
      7,
     </day>
     <month>
      September
     </month>
     <year>
      2025
     </year>
    </date>
    <date date-type="published">
     <day>
      12,
     </day>
     <month>
      September
     </month>
     <year>
      2025
     </year> 
    </date> 
    <date date-type="accepted">
     <day>
      12,
     </day>
     <month>
      October
     </month>
     <year>
      2025
     </year> 
    </date>
   </history>
   <permissions>
    <copyright-statement>
     © Copyright 2014 by authors and Scientific Research Publishing Inc. 
    </copyright-statement>
    <copyright-year>
     2014
    </copyright-year>
    <license>
     <license-p>
      This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/
     </license-p>
    </license>
   </permissions>
   <abstract>
    Predicting the progression from Mild Cognitive Impairment (MCI) to Alzheimer's Disease (AD) is a critical challenge for enabling early intervention and improving patient outcomes. While longitudinal multi-modal neuroimaging data holds immense potential for capturing the spatio-temporal dynamics of disease progression, its effective analysis is hampered by significant challenges: temporal heterogeneity (irregularly sampled scans), multi-modal misalignment, and the propensity of deep learning models to learn spurious, non-causal correlations. We propose CASCADE-Net, a novel end-to-end pipeline for robust and interpretable MCI-to-AD progression prediction. Our architecture introduces a Dynamic Temporal Alignment Module that employs a Neural Ordinary Differential Equation (Neural ODE) to model the continuous, underlying progression of pathology from irregularly sampled scans, effectively mapping heterogeneous patient data to a unified latent timeline. This aligned, noise-reduced spatio-temporal data is then processed by a predictive model featuring a novel Causal Spatial Attention mechanism. This mechanism not only identifies the critical brain regions and their evolution predictive of conversion but also incorporates a counterfactual constraint during training. This constraint ensures the learned features are causally linked to AD pathology by encouraging invariance to non-causal, confounder-based changes. Extensive experiments on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset demonstrate that CASCADE-Net significantly outperforms state-of-the-art sequential models in prognostic accuracy. Furthermore, our model provides highly interpretable, causally-grounded attention maps, offering valuable insights into the disease progression process and fostering greater clinical trust.
   </abstract>
   <kwd-group> 
    <kwd>
     Alzheimer’s Disease
    </kwd> 
    <kwd>
      Mild Cognitive Impairment
    </kwd> 
    <kwd>
      Prognosis
    </kwd> 
    <kwd>
      Neural ODE
    </kwd> 
    <kwd>
      Counterfactual Learning
    </kwd> 
    <kwd>
      Spatio-Temporal Modeling
    </kwd> 
    <kwd>
      Interpretable AI
    </kwd>
   </kwd-group>
  </article-meta>
 </front>
 <body>
  <sec id="s1">
   <title>1. Introduction</title>
   <p>Alzheimer’s Disease (AD) is a debilitating neurodegenerative disorder and the most common cause of dementia. The prodromal stage, known as Mild Cognitive Impairment (MCI), presents a critical window for intervention; however, not all individuals with MCI progress to AD. Accurately identifying which MCI patients are at the highest risk for conversion is therefore one of the most important challenges in modern neurology <xref ref-type="bibr" rid="scirp.146372-1">
     [1]
    </xref>.</p>
   <p>Longitudinal multi-modal neuroimaging, particularly Magnetic Resonance Imaging (MRI) and Positron Emission Tomography (PET), provides a powerful means to observe the in vivo evolution of AD pathology, including cortical atrophy, glucose hypometabolism, and amyloid-beta deposition <xref ref-type="bibr" rid="scirp.146372-2">
     [2]
    </xref>. Deep learning models, especially Recurrent Neural Networks (RNNs) and Long Short-Term Memory networks (LSTMs), have been applied to this sequential data for prognostic prediction <xref ref-type="bibr" rid="scirp.146372-3">
     [3]
    </xref>. Despite their promise, these approaches face three fundamental limitations:</p>
   <p>1) Temporal Heterogeneity: Patients are scanned at irregular, unpredictable intervals, violating the fixed-time-step assumption of standard RNNs. Simple interpolation or last-observation-carried-forward methods are inadequate for capturing complex non-linear disease dynamics.</p>
   <p>2) Multi-Modal Misalignment: Fusing information from different modalities (e.g., structural MRI with FDG-PET) is challenging due to different resolutions, contrasts, and the biological relationships between them.</p>
   <p>3) Spurious Correlations: Models may learn to rely on scanner-specific artifacts, demographic biases, or other confounding factors rather than the true biological signals of AD progression, leading to poor generalization and clinically untrustworthy predictions <xref ref-type="bibr" rid="scirp.146372-4">
     [4]
    </xref>.</p>
   <p>We propose CASCADE-Net (Causality-Aware Spatio-Temporal Dynamics Encoding Network), a novel architecture designed to overcome these hurdles. Our contributions are threefold:</p>
   <p>1) We introduce a Dynamic Temporal Alignment Module based on Neural ODEs to continuously model disease progression from irregularly sampled data, creating a regularized latent representation for each patient.</p>
   <p>2) We propose a Causal Spatial Attention mechanism that identifies critical spatio-temporal dynamics and is regularized by a counterfactual loss. This loss enforces causal invariance by ensuring model predictions are unchanged under perturbations that mimic confounding factors.</p>
   <p>3) We demonstrate through extensive experiments on the ADNI dataset that our end-to-end pipeline achieves state-of-the-art prognostic performance while providing interpretable, causally-justified attention maps that highlight the evolving pathological patterns indicative of AD conversion.</p>
  </sec><sec id="s2">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>2. Related Work</title>
   <sec id="s2_1">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>2.1. Prognostic Prediction in MCI</title>
    <p>Early machine learning approaches for MCI-to-AD conversion relied on hand-crafted features from single time-point images, such as cortical thickness or hippocampal volume <xref ref-type="bibr" rid="scirp.146372-5">
      [5]
     </xref>. With the advent of deep learning, convolutional neural networks (CNNs) were used to extract features from baseline scans <xref ref-type="bibr" rid="scirp.146372-6">
      [6]
     </xref>. However, these methods ignore the crucial temporal dimension. Subsequent work employed RNNs/LSTMs to model longitudinal data <xref ref-type="bibr" rid="scirp.146372-7">
      [7]
     </xref>. While effective on regularly sampled data, their performance degrades with real-world irregular sampling, a problem our method explicitly addresses.</p>
   </sec>
   <sec id="s2_2">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>2.2. Modeling Irregular Time Series</title>
    <p>To handle irregular sampling, methods like Phased LSTMs <xref ref-type="bibr" rid="scirp.146372-8">
      [8]
     </xref> and GRU-D <xref ref-type="bibr" rid="scirp.146372-9">
      [9]
     </xref> use time gates or decay mechanisms. More recently, Neural Ordinary Differential Equations (Neural ODEs) <xref ref-type="bibr" rid="scirp.146372-10">
      [10]
     </xref> have emerged as a powerful architecture for modeling continuous dynamics from discrete observations. They have shown promise in medical applications <xref ref-type="bibr" rid="scirp.146372-11">
      [11]
     </xref> but have not been fully explored for aligning multi-modal longitudinal neuroimaging data within a causal prediction architecture, which is our key innovation.</p>
   </sec>
   <sec id="s2_3">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>2.3. Interpretability and Causal Learning in Medical AI</title>
    <p>Attention mechanisms are widely used to interpret model decisions <xref ref-type="bibr" rid="scirp.146372-12">
      [12]
     </xref>. In neuroimaging, they help identify disease-relevant regions <xref ref-type="bibr" rid="scirp.146372-13">
      [13]
     </xref>. However, attention does not guarantee causality; it can highlight spurious correlations. Causal learning methods, particularly using counterfactual reasoning <xref ref-type="bibr" rid="scirp.146372-14">
      [14]
     </xref>, aim to mitigate this. Invariant risk minimization <xref ref-type="bibr" rid="scirp.146372-15">
      [15]
     </xref> and counterfactual augmentation <xref ref-type="bibr" rid="scirp.146372-16">
      [16]
     </xref> are relevant paradigms. Our counterfactual constraint draws inspiration from this line of work, applying it specifically to spatio-temporal attention in neuroimaging.</p>
   </sec>
  </sec><sec id="s3">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>3. The CASCADE-Net Architecture</title>
   <p>The overall architecture of CASCADE-Net is illustrated in <xref ref-type="fig" rid="fig1">
     Figure 1
    </xref>. The pipeline consists of two main components: 1) the Dynamic Temporal Alignment Module and 2) the Causal Spatio-Temporal Predictor.</p>
   <sec id="s3_1">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>3.1. Problem Formulation</title>
    <p>Let a patient’s longitudinal data be represented as a set of tuples:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi mathvariant="script">
         D 
       </mi> 
       <mo>
         = 
       </mo> 
       <msubsup> 
        <mrow> 
         <mrow> 
          <mo>
            { 
          </mo> 
          <mrow> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                t 
              </mi> 
              <mi>
                n 
              </mi> 
             </msub> 
             <mo>
               , 
             </mo> 
             <msubsup> 
              <mi>
                X 
              </mi> 
              <mi>
                n 
              </mi> 
              <mrow> 
               <mi>
                 M 
               </mi> 
               <mi>
                 R 
               </mi> 
               <mi>
                 I 
               </mi> 
              </mrow> 
             </msubsup> 
             <mo>
               , 
             </mo> 
             <msubsup> 
              <mi>
                X 
              </mi> 
              <mi>
                n 
              </mi> 
              <mrow> 
               <mi>
                 P 
               </mi> 
               <mi>
                 E 
               </mi> 
               <mi>
                 T 
               </mi> 
              </mrow> 
             </msubsup> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            } 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mi>
          N 
        </mi> 
       </msubsup> 
      </mrow> 
     </math>, where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math> is the time of the n-th visit relative to a baseline (e.g., 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>), and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          X 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math> are the corresponding multi-modal images. The visits are irregularly spaced, i.e., 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           + 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </msub> 
       <mo>
         − 
       </mo> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math> is not constant. The goal is to predict a binary label 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         Y 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math> indicating whether the patient will convert from MCI to AD within a predefined future time window (e.g., 3 years).</p>
    <fig id="fig1" position="float">
     <label>Figure 1</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 1. Overall architecture of the proposed CASCADE-Net pipeline. 1) Irregularly sampled multi-modal scans are processed by the Neural ODE-based Temporal Alignment Module to generate a regularly sampled latent trajectory. 2) This latent sequence is fed into a Spatio-Temporal Encoder (e.g., a 1D CNN + Transformer). 3) A Causal Spatial Attention module generates dynamic attention maps. 4) A counterfactual generator creates perturbed latent vectors based on the attention and a confounder model. The final prediction is made from the original latent sequence, and the model is trained with a combined task loss and counterfactual loss.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId29.jpeg?20251015114000" />
    </fig>
   </sec>
   <sec id="s3_2">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>3.2. Component 1: Dynamic Temporal Alignment Module</title>
    <p>This module encodes each snapshot 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            X 
          </mi> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mi>
             M 
           </mi> 
           <mi>
             R 
           </mi> 
           <mi>
             I 
           </mi> 
          </mrow> 
         </msubsup> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            X 
          </mi> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mi>
             P 
           </mi> 
           <mi>
             E 
           </mi> 
           <mi>
             T 
           </mi> 
          </mrow> 
         </msubsup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> into a latent vector 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           z 
         </mi> 
        </mstyle> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math> using a shared encoder network 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          E 
        </mi> 
        <mi>
          ϕ 
        </mi> 
       </msub> 
      </mrow> 
     </math>: This shared encoder 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          E 
        </mi> 
        <mi>
          ϕ 
        </mi> 
       </msub> 
      </mrow> 
     </math> is specifically designed to handle multi-modal misalignment. It first processes each modality through separate, modality-specific convolutional branches to extract features at the appropriate spatial scale and contrast for MRI and PET data respectively. The outputs of these branches are then concatenated and passed through subsequent shared convolutional layers. This design allows the network to first learn modality-specific representations before fusing them into a unified latent space, effectively aligning the heterogeneous information from different imaging protocols into a coherent representation for the subsequent Neural ODE processing.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           z 
         </mi> 
        </mstyle> 
        <mi>
          n 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          E 
        </mi> 
        <mi>
          ϕ 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msubsup> 
          <mi>
            X 
          </mi> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mi>
             M 
           </mi> 
           <mi>
             R 
           </mi> 
           <mi>
             I 
           </mi> 
          </mrow> 
         </msubsup> 
         <mo>
           , 
         </mo> 
         <msubsup> 
          <mi>
            X 
          </mi> 
          <mi>
            n 
          </mi> 
          <mrow> 
           <mi>
             P 
           </mi> 
           <mi>
             E 
           </mi> 
           <mi>
             T 
           </mi> 
          </mrow> 
         </msubsup> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (1)</p>
    <p>The sequence of latent vectors 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              t 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mstyle mathvariant="bold" mathsize="normal"> 
             <mi>
               z 
             </mi> 
            </mstyle> 
            <mn>
              1 
            </mn> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              t 
            </mi> 
            <mn>
              2 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mstyle mathvariant="bold" mathsize="normal"> 
             <mi>
               z 
             </mi> 
            </mstyle> 
            <mn>
              2 
            </mn> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              t 
            </mi> 
            <mi>
              N 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mstyle mathvariant="bold" mathsize="normal"> 
             <mi>
               z 
             </mi> 
            </mstyle> 
            <mi>
              N 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math> represents the patient’s state at observed time points.</p>
    <p>We model the continuous trajectory of the patient’s latent state using a Neural ODE. We define a neural network 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mi>
          θ 
        </mi> 
       </msub> 
      </mrow> 
     </math> that parameterizes the derivative of the latent state with respect to time:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mfrac> 
        <mrow> 
         <mtext>
           d 
         </mtext> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            z 
          </mi> 
         </mstyle> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mtext>
           d 
         </mtext> 
         <mi>
           t 
         </mi> 
        </mrow> 
       </mfrac> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          f 
        </mi> 
        <mi>
          θ 
        </mi> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            z 
          </mi> 
         </mstyle> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mi>
           t 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (2)</p>
    <p>Given an initial condition 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          z 
        </mi> 
       </mstyle> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            t 
          </mi> 
          <mn>
            0 
          </mn> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           z 
         </mi> 
        </mstyle> 
        <mn>
          0 
        </mn> 
       </msub> 
      </mrow> 
     </math>, the latent state at any time 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        t 
      </mi> 
     </math> can be found by solving the ODE:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          z 
        </mi> 
       </mstyle> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          t 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           z 
         </mi> 
        </mstyle> 
        <mn>
          0 
        </mn> 
       </msub> 
       <mo>
         + 
       </mo> 
       <mstyle displaystyle="true"> 
        <mrow> 
         <msubsup> 
          <mo>
            ∫ 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              t 
            </mi> 
            <mn>
              0 
            </mn> 
           </msub> 
          </mrow> 
          <mi>
            t 
          </mi> 
         </msubsup> 
         <mrow> 
          <msub> 
           <mi>
             f 
           </mi> 
           <mi>
             θ 
           </mi> 
          </msub> 
          <mrow> 
           <mo>
             ( 
           </mo> 
           <mrow> 
            <mstyle mathvariant="bold" mathsize="normal"> 
             <mi>
               z 
             </mi> 
            </mstyle> 
            <mrow> 
             <mo>
               ( 
             </mo> 
             <mi>
               τ 
             </mi> 
             <mo>
               ) 
             </mo> 
            </mrow> 
            <mo>
              , 
            </mo> 
            <mi>
              τ 
            </mi> 
           </mrow> 
           <mo>
             ) 
           </mo> 
          </mrow> 
          <mtext>
            d 
          </mtext> 
          <mi>
            τ 
          </mi> 
         </mrow> 
        </mrow> 
       </mstyle> 
      </mrow> 
     </math> (3)</p>
    <p>In practice, we use an ODE solver (e.g., Runge-Kutta) to perform this integration. This allows us to interpolate the latent state at any desired time.</p>
    <p>We define a regularized time grid 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi mathvariant="script">
         T 
       </mi> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          { 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            τ 
          </mi> 
          <mn>
            1 
          </mn> 
         </msub> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            τ 
          </mi> 
          <mn>
            2 
          </mn> 
         </msub> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            τ 
          </mi> 
          <mi>
            T 
          </mi> 
         </msub> 
        </mrow> 
        <mo>
          } 
        </mo> 
       </mrow> 
      </mrow> 
     </math> common to all patients (e.g., 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          τ 
        </mi> 
        <mi>
          i 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi>
         i 
       </mi> 
       <mo>
         × 
       </mo> 
       <mn>
         6 
       </mn> 
      </mrow> 
     </math> months). For each patient, we solve the Neural ODE to generate a aligned latent sequence:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           Z 
         </mi> 
        </mstyle> 
        <mrow> 
         <mtext>
           aligned 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mrow> 
        <mo>
          [ 
        </mo> 
        <mrow> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            z 
          </mi> 
         </mstyle> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              τ 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            z 
          </mi> 
         </mstyle> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              τ 
            </mi> 
            <mn>
              2 
            </mn> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           , 
         </mo> 
         <mo>
           ⋯ 
         </mo> 
         <mo>
           , 
         </mo> 
         <mstyle mathvariant="bold" mathsize="normal"> 
          <mi>
            z 
          </mi> 
         </mstyle> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              τ 
            </mi> 
            <mi>
              T 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ] 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (4)</p>
    <p>The parameters 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        ϕ 
      </mi> 
     </math> and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        θ 
      </mi> 
     </math> are learned end-to-end by minimizing a reconstruction loss (e.g., Mean Squared Error) between the observed latents 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           z 
         </mi> 
        </mstyle> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math> and the ODE solutions at the corresponding times 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          t 
        </mi> 
        <mi>
          n 
        </mi> 
       </msub> 
      </mrow> 
     </math>, ensuring the learned dynamics faithfully represent the patient’s true trajectory.</p>
   </sec>
   <sec id="s3_3">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>3.3. Component 2: Causal Spatio-Temporal Predictor</title>
    <p>The aligned latent sequence 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           Z 
         </mi> 
        </mstyle> 
        <mrow> 
         <mtext>
           aligned 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         ∈ 
       </mo> 
       <msup> 
        <mi>
          ℝ 
        </mi> 
        <mrow> 
         <mi>
           T 
         </mi> 
         <mo>
           × 
         </mo> 
         <mi>
           D 
         </mi> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is fed into the predictor.</p>
    <p>We use a 1D temporal convolutional network (TCN) <xref ref-type="bibr" rid="scirp.146372-17">
      [17]
     </xref> followed by a Transformer encoder <xref ref-type="bibr" rid="scirp.146372-12">
      [12]
     </xref> to capture complex long-range dependencies across time. This dual architecture of TCN followed by Transformer was chosen for their complementary strengths in capturing temporal dependencies. The TCN serves as an efficient local feature extractor, using its dilated convolutions to capture multi-scale, short-range patterns within the aligned latent sequence with minimal computational overhead. The Transformer encoder then processes these refined features to model global, long-range dependencies across the entire timeline through its self-attention mechanism. This allows every time point (e.g., baseline) to directly influence and be influenced by every other time point (e.g., 24-month), crucial for identifying complex, non-linear interactions between early and late-stage pathological changes that are characteristic of neurodegenerative progression.</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           H 
         </mi> 
        </mstyle> 
        <mrow> 
         <mtext>
           TCN 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mtext>
         TCN 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mstyle mathvariant="bold" mathsize="normal"> 
           <mi>
             Z 
           </mi> 
          </mstyle> 
          <mrow> 
           <mtext>
             aligned 
           </mtext> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (5)</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          H 
        </mi> 
       </mstyle> 
       <mo>
         = 
       </mo> 
       <mtext>
         Transformer 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mstyle mathvariant="bold" mathsize="normal"> 
           <mi>
             H 
           </mi> 
          </mstyle> 
          <mrow> 
           <mtext>
             TCN 
           </mtext> 
          </mrow> 
         </msub> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (6)</p>
    <p>The output 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          H 
        </mi> 
       </mstyle> 
       <mo>
         ∈ 
       </mo> 
       <msup> 
        <mi>
          ℝ 
        </mi> 
        <mrow> 
         <mi>
           T 
         </mi> 
         <mo>
           × 
         </mo> 
         <msup> 
          <mi>
            D 
          </mi> 
          <mo>
            ′ 
          </mo> 
         </msup> 
        </mrow> 
       </msup> 
      </mrow> 
     </math> is a refined spatio-temporal representation.</p>
    <p>An attention network 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          g 
        </mi> 
        <mi>
          ψ 
        </mi> 
       </msub> 
      </mrow> 
     </math> generates a dynamic attention weight for each feature dimension at each time step, effectively creating a spatio-temporal attention map 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         A 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           t 
         </mi> 
         <mo>
           , 
         </mo> 
         <mi>
           d 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         A 
       </mi> 
       <mo>
         = 
       </mo> 
       <mi>
         σ 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <msub> 
          <mi>
            g 
          </mi> 
          <mi>
            ψ 
          </mi> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mstyle mathvariant="bold" mathsize="normal"> 
           <mi>
             H 
           </mi> 
          </mstyle> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (7)</p>
    <p>where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        σ 
      </mi> 
     </math> is the sigmoid function, and 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         A 
       </mi> 
       <mo>
         ∈ 
       </mo> 
       <msup> 
        <mrow> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mrow> 
           <mn>
             0 
           </mn> 
           <mo>
             , 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
        <mrow> 
         <mi>
           T 
         </mi> 
         <mo>
           × 
         </mo> 
         <msup> 
          <mi>
            D 
          </mi> 
          <mo>
            ′ 
          </mo> 
         </msup> 
        </mrow> 
       </msup> 
      </mrow> 
     </math>. The attended representation is computed as:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathvariant="bold" mathsize="normal"> 
         <mi>
           H 
         </mi> 
        </mstyle> 
        <mrow> 
         <mtext>
           att 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi>
         A 
       </mi> 
       <mo>
         ⊙ 
       </mo> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          H 
        </mi> 
       </mstyle> 
      </mrow> 
     </math> (8)</p>
    <p>This attended representation is then global-average-pooled over time and passed through a final classifier (a linear layer) to produce the prediction probability 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mi>
          Y 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         = 
       </mo> 
       <mi>
         P 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           Y 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           | 
         </mo> 
         <mi mathvariant="script">
           D 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>.</p>
    <p>This is the core of our causal reasoning. We define a confounder distribution 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          c 
        </mi> 
       </msub> 
      </mrow> 
     </math>, which represents non-causal changes. For neuroimaging, this could be a model of healthy aging or scanner drift. For a given patient’s latent vector at time 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        τ 
      </mi> 
     </math>, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathvariant="bold" mathsize="normal"> 
        <mi>
          z 
        </mi> 
       </mstyle> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          τ 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, and its attention 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         A 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          τ 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, we generate a counterfactual latent vector:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mstyle mathsize="normal" mathvariant="bold"> 
        <mover accent="true"> 
         <mi>
           z 
         </mi> 
         <mo>
           ˜ 
         </mo> 
        </mover> 
       </mstyle> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          τ 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         = 
       </mo> 
       <mstyle mathsize="normal" mathvariant="bold"> 
        <mi>
          z 
        </mi> 
       </mstyle> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          τ 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ⊙ 
       </mo> 
       <mi>
         A 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mi>
          τ 
        </mi> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
         where 
       </mtext> 
       <mtext>
           
       </mtext> 
       <mtext>
           
       </mtext> 
       <mi>
         ϵ 
       </mi> 
       <mo>
         ~ 
       </mo> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          c 
        </mi> 
       </msub> 
      </mrow> 
     </math> (9)</p>
    <p>The confounder distribution 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          c 
        </mi> 
       </msub> 
      </mrow> 
     </math> is central to the validity of our counterfactual constraint. In this work, we model 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          c 
        </mi> 
       </msub> 
      </mrow> 
     </math> as a zero-mean, isotropic Gaussian distribution, 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          c 
        </mi> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mi mathvariant="script">
         N 
       </mi> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mn>
           0 
         </mn> 
         <mo>
           , 
         </mo> 
         <msup> 
          <mi>
            σ 
          </mi> 
          <mn>
            2 
          </mn> 
         </msup> 
         <mi>
           I 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math>, where 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        I 
      </mi> 
     </math> is the identity matrix. This simple prior is chosen to represent non-informative, non-causal variations, such as minor scanner noise or benign anatomical differences not linked to AD pathology. The standard deviation 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        σ 
      </mi> 
     </math> is a key hyperparameter that controls the magnitude of the counterfactual perturbation. Its value was set to 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         σ 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         0.15 
       </mn> 
      </mrow> 
     </math> through a grid search on the validation set to maximize prognostic performance; this value was found to generate meaningful counterfactuals that altered the latent representation without pushing it into implausible regions of the feature space. While effective, the simplicity of this prior is a limitation, and learning a more sophisticated, data-driven confounder model from large-scale control populations is a focus of future work.</p>
    <p>This perturbation applies a confounder-based change specifically to the features the model found most salient. The counterfactual latent sequence 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mstyle mathsize="normal" mathvariant="bold"> 
         <mover accent="true"> 
          <mi>
            Z 
          </mi> 
          <mo>
            ˜ 
          </mo> 
         </mover> 
        </mstyle> 
        <mrow> 
         <mtext>
           aligned 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is processed by the same predictor to get a counterfactual prediction 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
       <mi>
         Y 
       </mi> 
       <mo>
         ˜ 
       </mo> 
      </mover> 
     </math>.</p>
    <p>The model is trained with a combined loss function:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         ℒ 
       </mi> 
       <mo>
         = 
       </mo> 
       <msub> 
        <mi>
          ℒ 
        </mi> 
        <mrow> 
         <mtext>
           task 
         </mtext> 
        </mrow> 
       </msub> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mover accent="true"> 
          <mi>
            Y 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
         <mo>
           , 
         </mo> 
         <mi>
           Y 
         </mi> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
       <mo>
         + 
       </mo> 
       <mi>
         β 
       </mi> 
       <mo>
         ⋅ 
       </mo> 
       <msub> 
        <mi>
          ℒ 
        </mi> 
        <mrow> 
         <mtext>
           CF 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> (10)</p>
    <p>The task loss 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          ℒ 
        </mi> 
        <mrow> 
         <mtext>
           task 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is standard binary cross-entropy. The counterfactual loss 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          ℒ 
        </mi> 
        <mrow> 
         <mtext>
           CF 
         </mtext> 
        </mrow> 
       </msub> 
      </mrow> 
     </math> is the Kullback-Leibler (KL) divergence between the original and counterfactual predictions:</p>
    <p>
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          ℒ 
        </mi> 
        <mrow> 
         <mtext>
           CF 
         </mtext> 
        </mrow> 
       </msub> 
       <mo>
         = 
       </mo> 
       <mtext>
         KL 
       </mtext> 
       <mrow> 
        <mo>
          ( 
        </mo> 
        <mrow> 
         <mi>
           P 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              Y 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             | 
           </mo> 
           <mstyle mathsize="normal" mathvariant="bold"> 
            <mi>
              Z 
            </mi> 
           </mstyle> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           ∥ 
         </mo> 
         <mi>
           P 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              Y 
            </mi> 
            <mo>
              ˜ 
            </mo> 
           </mover> 
           <mo>
             | 
           </mo> 
           <mstyle mathsize="normal" mathvariant="bold"> 
            <mover accent="true"> 
             <mi>
               Z 
             </mi> 
             <mo>
               ˜ 
             </mo> 
            </mover> 
           </mstyle> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
        <mo>
          ) 
        </mo> 
       </mrow> 
      </mrow> 
     </math> (11)</p>
    <p>This loss penalizes the model if its prediction changes under this confounder-based perturbation of attended features. It forces the predictor to rely on features whose predictive power is invariant to these confounders, i.e., features that are more likely to be causally linked to AD pathology. The hyperparameter 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
        β 
      </mi> 
     </math> controls the strength of this causal constraint.</p>
   </sec>
  </sec><sec id="s4">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>4. Experiments and Results</title>
   <sec id="s4_1">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>Dataset and Experimental Setup</title>
    <p>We used data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) <xref ref-type="bibr" rid="scirp.146372-18">
      [18]
     </xref>. Our cohort consisted of 500 MCI patients (250 converters (MCI-C), 250 non-converters (MCI-NC)) with at least three longitudinal T1-weighted MRI and FDG-PET scans each. Data was preprocessed: MRIs were segmented and normalized to a common space; PET scans were co-registered to their corresponding MRI and intensity normalized. We used data from baseline, 12-month, and 24-month visits as inputs and defined conversion as a clinical diagnosis of AD within 36 months from baseline.</p>
    <p>We implemented CASCADE-Net in PyTorch <xref ref-type="bibr" rid="scirp.146372-19">
      [19]
     </xref> using the TorchDiffEq package. The model was trained with the Adam optimizer <xref ref-type="bibr" rid="scirp.146372-20">
      [20]
     </xref> for 100 epochs with an initial learning rate of 1e<sup>−</sup><sup>4</sup>. We used a 5-fold cross-validation strategy. We compared against several strong baselines:</p>
    <p>Performance was evaluated using Area Under the ROC Curve (AUC), Accuracy (ACC), Sensitivity (SEN), and Specificity (SPEC).</p>
   </sec>
  </sec><sec id="s5">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>5. Training Dynamics and Convergence Analysis</title>
   <p>The training dynamics and convergence behavior of CASCADE-Net, compared against state-of-the-art baseline models, provide crucial insights into the optimization efficiency and stability of our proposed architecture. <xref ref-type="fig" rid="fig2">
     Figure 2
    </xref> (training/validation loss curves) and <xref ref-type="fig" rid="fig3">
     Figure 3
    </xref> (convergence speed comparison) collectively demonstrate several key advantages of our causality-aware spatio-temporal framework.</p>
   <sec id="s5_1">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>5.1. Superior Convergence Characteristics</title>
    <p>CASCADE-Net exhibited significantly faster convergence compared to all baseline models, stabilizing after approximately 25 epochs—2.4× faster than the standard LSTM (60 epochs) and 1.4× faster than the Neural ODE approach (35 epochs). This accelerated convergence can be attributed to several architectural innovations:</p>
    <p>Efficient Gradient Propagation: The integration of Neural ODE solvers with causal attention mechanisms enables more stable gradient flow during backpropagation. Unlike traditional recurrent architectures that suffer from vanishing gradient problems in long temporal sequences, our continuous-time formulation maintains gradient integrity across irregular time intervals.</p>
    <p>Regularized Learning Dynamics: The causality-aware constraints act as an implicit regularizer, preventing the model from overfitting to spurious correlations in the early training phases. This regularization effect is particularly evident in the smooth, monotonic decrease of both training and validation losses in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref>, contrasting with the occasional fluctuations observed in baseline models.</p>
    <fig id="fig2" position="float">
     <label>Figure 2</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 2. Training and validation loss curves for CASCADE-Net and baseline models. CASCADE-Net demonstrates faster convergence and lower final loss values compared to all baseline approaches. The dashed vertical lines indicate convergence points where each model’s loss stabilized within 1% of its final value.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId124.jpeg?20251015114004" />
    </fig>
    <fig id="fig3" position="float">
     <label>Figure 3</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 3. Convergence speed comparison showing the number of epochs required for each model to reach stable performance. CASCADE-Net converges 2.4× faster than standard LSTM and 1.4× faster than Neural ODE approaches.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId125.jpeg?20251015114005" />
    </fig>
   </sec>
   <sec id="s5_2">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>5.2. Enhanced Optimization Stability</title>
    <p>The loss curves in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> reveal distinct optimization patterns across different architectures:</p>
    <p>CASCADE-Net demonstrated the most stable optimization trajectory, with both training and validation losses decreasing smoothly without significant oscillations. This stability suggests that the spatio-temporal alignment module effectively resolves the distribution mismatches between irregularly sampled inputs and the regularized latent space.</p>
    <p>Neural ODE-based approaches showed improved stability over traditional RNN variants but still exhibited minor fluctuations, particularly during the first 20 epochs. This indicates that while continuous-time modeling helps, it requires additional architectural components (as implemented in CASCADE-Net) to achieve optimal stability.</p>
    <p>GRU-D and Standard LSTM models displayed pronounced oscillations throughout training, reflecting the challenges of handling irregular sampling patterns and long-term dependencies without explicit temporal alignment mechanisms.</p>
   </sec>
   <sec id="s5_3">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>5.3. Generalization Performance</title>
    <p>The validation loss curves in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> provide compelling evidence for CASCADE-Net’s superior generalization capabilities:</p>
    <p>Minimal Overfitting Gap: The small divergence between training and validation losses ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mn>
         0.0085 
       </mn> 
      </mrow> 
     </math>) indicates excellent generalization, significantly outperforming Neural ODE ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mn>
         0.012 
       </mn> 
      </mrow> 
     </math>), GRU-D ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mn>
         0.015 
       </mn> 
      </mrow> 
     </math>), and LSTM ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mtext>
         Δ 
       </mtext> 
       <mo>
         = 
       </mo> 
       <mn>
         0.018 
       </mn> 
      </mrow> 
     </math>) models. This reduced generalization gap demonstrates the effectiveness of our causal regularization in preventing overfitting to training-specific patterns.</p>
    <p>Early Stopping Robustness: CASCADE-Net reached near-optimal performance much earlier than comparative models, as clearly shown in <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref>, suggesting practical advantages for clinical applications where computational resources may be limited. The model maintained stable performance after convergence, without exhibiting the performance degradation observed in some baseline models during extended training.</p>
   </sec>
  </sec><sec id="s6">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>6. Comprehensive Performance Analysis</title>
   <p>This section presents a comprehensive evaluation of CASCADE-Net’s predictive capabilities for MCI-to-AD conversion prediction, encompassing ROC analysis, confusion matrix examination, and detailed classification metrics compared to established baseline models.</p>
   <sec id="s6_1">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.1. Receiver Operating Characteristic Analysis</title>
    <p>The Receiver Operating Characteristic analysis in <xref ref-type="fig" rid="fig4">
      Figure 4
     </xref> provides critical insights into the discriminatory power of CASCADE-Net. The model achieves an exceptional AUC value of 0.87 as in <xref ref-type="fig" rid="fig5">
      Figure 5
     </xref> and <xref ref-type="fig" rid="fig6">
      Figure 6
     </xref>, significantly outperforming all comparative models.</p>
    <p>Clinical Interpretation of AUC Performance:</p>
    <fig id="fig4" position="float">
     <label>Figure 4</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 4. Comparative ROC curves demonstrating superior performance of CASCADE-Net across all baseline models. The analysis reveals CASCADE-Net’s enhanced discriminatory power with an AUC of 0.87.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId134.jpeg?20251015114007" />
    </fig>
    <fig id="fig5" position="float">
     <label>Figure 5</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 5. Detailed ROC curve for CASCADE-Net showing excellent discriminatory performance (AUC = 0.87) with narrow confidence intervals, indicating robust model performance.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId135.jpeg?20251015114007" />
    </fig>
    <fig id="fig6" position="float">
     <label>Figure 6</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 6. AUC performance comparison between baseline models and CASCADE-Net, showing progressive improvement from basic to advanced architectures.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId136.jpeg?20251015114007" />
    </fig>
    <p>The Precision-Recall analysis reveals CASCADE-Net’s strong clinical applicability, maintaining high precision (AP = 0.85) across all recall levels. This consistency is particularly valuable for clinical deployment where different operating points may be required based on specific clinical scenarios.</p>
   </sec>
   <sec id="s6_2">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.2. Performance Metrics Interpretation</title>
    <p>The quantitative results demonstrate CASCADE-Net’s superior performance across all evaluation metrics:</p>
    <p>Area Under the Curve (AUC): 0.87 ± 0.01</p>
    <p>CASCADE-Net achieves outstanding discriminatory power, representing a 19% relative improvement over Baseline CNN and 13% improvement over Neural ODE ablation. This indicates excellent ability to distinguish between converters and non-converters across all classification thresholds.</p>
    <p>Accuracy: 79.4% ± 1.2%</p>
    <p>The model correctly classifies approximately 4 out of 5 patients, representing substantial 16% and 11% improvements over Baseline CNN and Standard LSTM respectively, demonstrating effective temporal alignment and causal attention mechanisms.</p>
    <p>Balanced Sensitivity (0.80) and Specificity (0.79)</p>
    <p>The nearly identical values indicate excellent performance balance without significant class bias. The 80% sensitivity enables early intervention for true converters, while 79% specificity reduces unnecessary interventions for non-converters.</p>
   </sec>
   <sec id="s6_3">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.3. Confusion Matrix Analysis</title>
    <p>The confusion matrix analysis in <xref ref-type="fig" rid="fig7">
      Figure 7
     </xref> reveals several key advantages:</p>
    <p>Reduced False Positives: CASCADE-Net demonstrates significantly fewer false positives (21% vs 34% for baseline), reducing unnecessary treatments and patient anxiety in clinical settings.</p>
    <p>Enhanced True Positive Detection: The 80% true positive rate provides longer intervention windows for actual converters, enabling more effective treatment planning.</p>
    <p>Balanced Classification: Nearly equal sensitivity and specificity values indicate excellent balance between identifying true converters and avoiding false alarms, particularly valuable given typical class imbalances in clinical datasets.</p>
    <fig id="fig7" position="float">
     <label>Figure 7</label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 7. Comparative confusion matrices demonstrating CASCADE-Net’s superior performance versus baseline CNN. The analysis reveals reduced false positives and improved true positive detection.</title>
     </caption>
     <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId137.jpeg?20251015114008" />
    </fig>
   </sec>
   <sec id="s6_4">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.4. Detailed Classification Performance</title>
    <p>The classification reports reveal CASCADE-Net’s superior characteristics:</p>
    <p>Precision: Substantially higher for both classes (0.786 vs 0.652 for MCI-NC; 0.790 vs 0.678 for MCI-C), demonstrating more reliable positive predictions.</p>
    <p>Recall: Balanced across both classes (0.792 vs 0.784), unlike the baseline which shows significant imbalance (0.704 vs 0.624).</p>
    <p>F1-Score: Markedly higher values (0.789 vs 0.677 for MCI-NC; 0.787 vs 0.650 for MCI-C), representing 16.5% and 21.1% improvements respectively. This statistics are well summarised in <xref ref-type="table" rid="table1">
      Table 1
     </xref> and <xref ref-type="table" rid="table2">
      Table 2
     </xref> for CASCADE Net and Baseline CNN respectively.</p>
    <table-wrap id="table1">
     <label>
      <xref ref-type="table" rid="table1">
       Table 1
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Table 1. Detailed classification report for CASCADE-Net.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="23.08%"><p style="text-align:center">Class</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.22%"><p style="text-align:center">Precision</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.23%"><p style="text-align:center">Recall</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.23%"><p style="text-align:center">F1-Score</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.23%"><p style="text-align:center">Support</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="23.08%"><p style="text-align:center">MCI-NC</p></td> 
       <td class="custom-top-td acenter" width="19.22%"><p style="text-align:center">0.786</p></td> 
       <td class="custom-top-td acenter" width="19.23%"><p style="text-align:center">0.792</p></td> 
       <td class="custom-top-td acenter" width="19.23%"><p style="text-align:center">0.789</p></td> 
       <td class="custom-top-td acenter" width="19.23%"><p style="text-align:center">125</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="23.08%"><p style="text-align:center">MCI-C</p></td> 
       <td class="acenter" width="19.22%"><p style="text-align:center">0.790</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.784</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.787</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">125</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="23.08%"><p style="text-align:center">Accuracy</p></td> 
       <td class="acenter" width="76.92%" colspan="4"><p style="text-align:center">0.788</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="23.08%"><p style="text-align:center">Macro Avg</p></td> 
       <td class="acenter" width="19.22%"><p style="text-align:center">0.788</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.788</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.788</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">250</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="23.08%"><p style="text-align:center">Weighted Avg</p></td> 
       <td class="custom-bottom-td acenter" width="19.22%"><p style="text-align:center">0.788</p></td> 
       <td class="custom-bottom-td acenter" width="19.23%"><p style="text-align:center">0.788</p></td> 
       <td class="custom-bottom-td acenter" width="19.23%"><p style="text-align:center">0.788</p></td> 
       <td class="custom-bottom-td acenter" width="19.23%"><p style="text-align:center">250</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table2">
     <label>
      <xref ref-type="table" rid="table2">
       Table 2
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Table 2. Detailed classification report for Baseline CNN.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="23.08%"><p style="text-align:center">Class</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.22%"><p style="text-align:center">Precision</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.23%"><p style="text-align:center">Recall</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.23%"><p style="text-align:center">F1-Score</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="19.23%"><p style="text-align:center">Support</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="23.08%"><p style="text-align:center">MCI-NC</p></td> 
       <td class="custom-top-td acenter" width="19.22%"><p style="text-align:center">0.652</p></td> 
       <td class="custom-top-td acenter" width="19.23%"><p style="text-align:center">0.704</p></td> 
       <td class="custom-top-td acenter" width="19.23%"><p style="text-align:center">0.677</p></td> 
       <td class="custom-top-td acenter" width="19.23%"><p style="text-align:center">125</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="23.08%"><p style="text-align:center">MCI-C</p></td> 
       <td class="acenter" width="19.22%"><p style="text-align:center">0.678</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.624</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.650</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">125</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="23.08%"><p style="text-align:center">Accuracy</p></td> 
       <td class="acenter" width="76.92%" colspan="4"><p style="text-align:center">0.664</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="23.08%"><p style="text-align:center">Macro Avg</p></td> 
       <td class="acenter" width="19.22%"><p style="text-align:center">0.665</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.664</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">0.663</p></td> 
       <td class="acenter" width="19.23%"><p style="text-align:center">250</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="23.08%"><p style="text-align:center">Weighted Avg</p></td> 
       <td class="custom-bottom-td acenter" width="19.22%"><p style="text-align:center">0.665</p></td> 
       <td class="custom-bottom-td acenter" width="19.23%"><p style="text-align:center">0.664</p></td> 
       <td class="custom-bottom-td acenter" width="19.23%"><p style="text-align:center">0.663</p></td> 
       <td class="custom-bottom-td acenter" width="19.23%"><p style="text-align:center">250</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <p>CASCADE-Net exhibits exceptional balance across performance metrics:</p>
    <p>Class Balance: Minimal differences between classes (0.006 across metrics) compared to baseline’s 12.8% recall difference.</p>
    <p>Metric Consistency: Close alignment between precision, recall, and F1-score indicates robust and reliable classification behavior.</p>
   </sec>
   <sec id="s6_5">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.5. Clinical Implications and Significance</title>
    <p>The performance characteristics have profound clinical implications:</p>
    <p>Early Intervention: 80% sensitivity enables identification of most future AD cases, facilitating earlier interventions.</p>
    <p>Resource Optimization: 79% specificity reduces unnecessary referrals and testing, optimizing healthcare resource allocation.</p>
    <p>Trustworthy Predictions: Balanced performance across metrics ensures clinically reliable predictions without systematic bias.</p>
    <p>Computational Efficiency: Rapid convergence and stable optimization make CASCADE-Net feasible for deployment in resource-constrained healthcare settings.</p>
   </sec>
   <sec id="s6_6">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.6. Statistical Significance and Comparative Advantage</title>
    <p>The performance improvements are statistically significant (p &lt; 0.01) across all metrics. Small standard deviations (±0.01 for AUC, ±1.2% for accuracy) indicate consistent performance across validation folds, demonstrating robustness and generalizability.</p>
    <p>The ablation study reveals progressive improvements:</p>
   </sec>
   <sec id="s6_7">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.7. Conclusion</title>
    <p>The comprehensive performance analysis demonstrates that CASCADE-Net achieves state-of-the-art predictive performance while providing balanced, clinically-reliable predictions. The model’s excellent discriminatory power (AUC: 0.87), balanced sensitivity/specificity (0.80/0.79), and consistent performance across metrics position it as a valuable tool for MCI-to-AD conversion prediction with significant potential impact on patient care and resource allocation in clinical practice.</p>
   </sec>
   <sec id="s6_8">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.8. Theoretical Insights</title>
    <p>The convergence behavior aligns with our theoretical framework regarding causality-aware learning:</p>
    <p>The rapid stabilization of loss curves in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> supports our hypothesis that explicit modeling of causal relationships reduces the hypothesis space that the model needs to explore during training. By incorporating domain knowledge about temporal dependencies and causal mechanisms in neurodegenerative progression, CASCADE-Net avoids learning spurious correlations that often plague purely data-driven approaches.</p>
    <p>Furthermore, the parallel decrease in both training and validation losses suggests that the causal inductive biases embedded in our architecture align well with the underlying data-generating process of MCI-to-AD conversion, validating our approach to integrating clinical domain knowledge with deep learning methodologies.</p>
    <p>In conclusion, the training dynamics and convergence analysis in <xref ref-type="fig" rid="fig2">
      Figure 2
     </xref> and <xref ref-type="fig" rid="fig3">
      Figure 3
     </xref> not only demonstrate CASCADE-Net’s practical advantages in terms of efficiency and stability but also provide empirical validation of our theoretical framework for causality-aware spatio-temporal modeling in clinical prognostic prediction.</p>
   </sec>
   <sec id="s6_9">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.9. Prognostic Performance Comparison</title>
    <p>As shown in <xref ref-type="table" rid="table3">
      Table 3
     </xref>, CASCADE-Net achieves the best performance across all metrics, significantly outperforming all baseline models (paired t-test, p &lt; 0.01). This demonstrates the overall effectiveness of our integrated approach. The improvement over the Standard LSTM highlights the benefit of handling irregular sampling. The gain over GRU-D suggests the Neural ODE offers a more powerful continuous dynamic model. The significant jump over the Neural ODE + Classifier ablation underscores the critical contribution of our novel Causal Spatial Attention mechanism.</p>
   </sec>
   <sec id="s6_10">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.10. Ablation Study</title>
    <p>
     <xref ref-type="table" rid="table4">
      Table 4
     </xref> presents the results of our ablation study. Each component contributes positively to the final performance. Removing the Temporal Alignment module causes the largest performance drop, confirming its necessity. Using a standard attention mechanism without the counterfactual loss ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mi>
         β 
       </mi> 
       <mo>
         = 
       </mo> 
       <mn>
         0 
       </mn> 
      </mrow> 
     </math>) leads to a noticeable drop in AUC, suggesting that without the causal constraint, the model learns slightly less robust features. Replacing the counterfactual loss with a standard 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          L 
        </mi> 
        <mn>
          1 
        </mn> 
       </msub> 
      </mrow> 
     </math> sparsity constraint on attention performs worse, indicating that our CF loss does more than just sparsify attention—it guides it toward causal features.</p>
    <table-wrap id="table3">
     <label>
      <xref ref-type="table" rid="table3">
       Table 3
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Table 3. Performance comparison of different models for MCI-to-AD conversion prediction.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="29.50%"><p style="text-align:center">Model</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="17.62%"><p style="text-align:center">AUC</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="17.62%"><p style="text-align:center">ACC</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="17.62%"><p style="text-align:center">SEN</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="17.64%"><p style="text-align:center">SPEC</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td acenter" width="29.50%"><p style="text-align:center">Baseline CNN</p></td> 
       <td class="custom-top-td acenter" width="17.62%"><p style="text-align:center">0.73 ± 0.03</p></td> 
       <td class="custom-top-td acenter" width="17.62%"><p style="text-align:center">68.2 ± 2.1</p></td> 
       <td class="custom-top-td acenter" width="17.62%"><p style="text-align:center">0.70</p></td> 
       <td class="custom-top-td acenter" width="17.64%"><p style="text-align:center">0.66</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="29.50%"><p style="text-align:center">Standard LSTM</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">0.77 ± 0.02</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">71.5 ± 1.8</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">0.72</p></td> 
       <td class="acenter" width="17.64%"><p style="text-align:center">0.71</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="29.50%"><p style="text-align:center">GRU-D</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">0.80 ± 0.02</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">73.8 ± 1.5</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">0.75</p></td> 
       <td class="acenter" width="17.64%"><p style="text-align:center">0.73</p></td> 
      </tr> 
      <tr> 
       <td class="acenter" width="29.50%"><p style="text-align:center">Neural ODE + Classifier</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">0.82 ± 0.02</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">75.1 ± 1.6</p></td> 
       <td class="acenter" width="17.62%"><p style="text-align:center">0.76</p></td> 
       <td class="acenter" width="17.64%"><p style="text-align:center">0.74</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td acenter" width="29.50%"><p style="text-align:center">CASCADE-Net (Ours)</p></td> 
       <td class="custom-bottom-td acenter" width="17.62%"><p style="text-align:center">0.87 ± 0.01</p></td> 
       <td class="custom-bottom-td acenter" width="17.62%"><p style="text-align:center">79.4 ± 1.2</p></td> 
       <td class="custom-bottom-td acenter" width="17.62%"><p style="text-align:center">0.80</p></td> 
       <td class="custom-bottom-td acenter" width="17.64%"><p style="text-align:center">0.79</p></td> 
      </tr> 
     </table>
    </table-wrap>
    <table-wrap id="table4">
     <label>
      <xref ref-type="table" rid="table4">
       Table 4
      </xref></label>
     <caption>
      <title>
       <xref ref-type="bibr" rid="scirp.146372-"></xref>Table 4. Ablation study on the components of CASCADE-Net.</title>
     </caption>
     <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
      <tr> 
       <td class="custom-bottom-td custom-top-td acenter" width="72.51%"><p style="text-align:center">Model Variant</p></td> 
       <td class="custom-bottom-td custom-top-td acenter" width="27.49%"><p style="text-align:center">AUC</p></td> 
      </tr> 
      <tr> 
       <td class="custom-top-td aleft" width="72.51%"><p style="text-align:left">Full CASCADE-Net</p></td> 
       <td class="custom-top-td acenter" width="27.49%"><p style="text-align:center">0.87</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="72.51%"><p style="text-align:left">- without Temporal Alignment (use GRU-D instead)</p></td> 
       <td class="acenter" width="27.49%"><p style="text-align:center">0.82</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="72.51%"><p style="text-align:left">- without Attention (use average pooling)</p></td> 
       <td class="acenter" width="27.49%"><p style="text-align:center">0.83</p></td> 
      </tr> 
      <tr> 
       <td class="aleft" width="72.51%"><p style="text-align:left">- without Counterfactual Loss ( 
         <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
           <mi>
             β 
           </mi> 
           <mo>
             = 
           </mo> 
           <mn>
             0 
           </mn> 
          </mrow> 
         </math>)</p></td> 
       <td class="acenter" width="27.49%"><p style="text-align:center">0.84</p></td> 
      </tr> 
      <tr> 
       <td class="custom-bottom-td aleft" width="72.51%"><p style="text-align:left">- with Standard ( 
         <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
           <msub> 
            <mi>
              L 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
          </mrow> 
         </math>) Attention Regularization</p></td> 
       <td class="custom-bottom-td acenter" width="27.49%"><p style="text-align:center">0.85</p></td> 
      </tr> 
     </table>
    </table-wrap>
   </sec>
   <sec id="s6_11">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>6.11. Interpretability and Qualitative Analysis</title>
    <p>
     <xref ref-type="fig" rid="fig8">
      Figure 8
     </xref> visualizes the dynamic attention maps for example patients. For a converter patient, the attention progressively focuses on regions known to be severely affected in AD, such as the medial temporal lobe (including the hippocampus), entorhinal cortex, and temporoparietal association areas. This evolving pattern aligns with the known Braak staging of neurofibrillary tangle pathology <xref ref-type="bibr" rid="scirp.146372-21">
      [21]
     </xref>. For a non-converter, the attention is either more diffuse, stable, or focuses on areas less specific to AD progression. This interpretable output provides a clear, data-driven rationale for the model’s prediction, which can be invaluable for clinical experts.</p>
   </sec>
  </sec><sec id="s7">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>7. CASCADE-Net Algorithm</title>
   <p>The CASCADE-Net procedure, formalized in Algorithm 1, integrates our core contributions into an end-to-end learning framework. The algorithm begins by processing a patient’s irregularly sampled, multi-modal scans through the Dynamic Temporal Alignment Module, which leverages a Neural ODE to map the heterogeneous inputs onto a unified latent timeline. This aligned representation is then passed to the Causal Spatio-Temporal Predictor, where a Transformer encoder refined by a novel Causal Spatial Attention mechanism identifies critical dynamic patterns. Crucially, the attended features are subjected to a counterfactual regularization step that perturbs them based on a confounder model, and the resulting counterfactual loss ensures the model’s predictions remain invariant to non-causal variations, thereby promoting robust and causally-grounded feature learning. The entire pipeline is optimized with a combined objective of accurate prognosis and causal invariance. For making predictions on new data, the streamlined inference process is detailed in Algorithm 2.</p>
   <fig id="fig8" position="float">
    <label>Figure 8</label>
    <caption>
     <title>
      <xref ref-type="bibr" rid="scirp.146372-"></xref>Figure 8. Visualization of learned dynamic attention maps for a converter (MCI-C) and a non-converter (MCI-NC) patient over the regularized time grid. Warmer colors indicate higher attention weights. The MCI-C patient shows increasing attention in known AD-vulnerable regions like the hippocampus and temporal cortex over time, while the MCI-NC patient shows a more stable or diffuse pattern.</title>
    </caption>
    <graphic mimetype="image" position="float" xlink:type="simple" xlink:href="https://html.scirp.org/file/9601737-rId146.jpeg?20251015114013" />
   </fig>
   <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
    <tr> 
     <td class="custom-bottom-td custom-top-td aleft" width="100.00%"><p style="text-align:left">Algorithm 1 CASCADE-Net Training Procedure</p></td> 
    </tr> 
    <tr> 
     <td class="custom-top-td aleft" width="100.00%"><p style="text-align:left">Input: Dataset 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi mathvariant="script">
           D 
         </mi> 
         <mo>
           = 
         </mo> 
         <msubsup> 
          <mrow> 
           <mrow> 
            <mo>
              { 
            </mo> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi mathvariant="script">
                  D 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
               <mo>
                 , 
               </mo> 
               <msub> 
                <mi>
                  Y 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mo>
              } 
            </mo> 
           </mrow> 
          </mrow> 
          <mrow> 
           <mi>
             i 
           </mi> 
           <mo>
             = 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mi>
            M 
          </mi> 
         </msubsup> 
        </mrow> 
       </math>, Confounder distribution 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            p 
          </mi> 
          <mi>
            c 
          </mi> 
         </msub> 
        </mrow> 
       </math>, Hyperparameters: 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          β 
        </mi> 
       </math>, 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mrow> 
          <mo>
            { 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              τ 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mn>
             ... 
           </mn> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mi>
              τ 
            </mi> 
            <mi>
              T 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            } 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">Output: Trained parameters 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mtext>
           Θ 
         </mtext> 
         <mo>
           = 
         </mo> 
         <mrow> 
          <mo>
            { 
          </mo> 
          <mrow> 
           <mi>
             ϕ 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             θ 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             ψ 
           </mi> 
           <mo>
             , 
           </mo> 
           <mi>
             ω 
           </mi> 
          </mrow> 
          <mo>
            } 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">Summary: 1) Align timelines to a grid via a Neural ODE. </p><p style="text-align:left">2) Predict diagnosis with a Transformer-TCN and counterfactual regularizer.</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">1) Initialize model parameters 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
          Θ 
        </mtext> 
       </math>.</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">2) For epoch 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
        </mrow> 
       </math> to 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            N 
          </mi> 
          <mrow> 
           <mtext>
             epochs 
           </mtext> 
          </mrow> 
         </msub> 
        </mrow> 
       </math>:</p></td> 
    </tr> 
    <tr> 
     <td class="custom-bottom-td aleft" width="100.00%"><p style="text-align:left">a) For each mini-batch 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           ℬ 
         </mi> 
         <mo>
           ⊂ 
         </mo> 
         <mi mathvariant="script">
           D 
         </mi> 
        </mrow> 
       </math>:</p></td> 
    </tr> 
   </table>
   <p>Continued</p>
   <table class="MsoTableGrid custom-table" border="0" cellspacing="0" cellpadding="0"> 
    <tr> 
     <td class="custom-top-td aleft" width="100.00%"><p style="text-align:left">i) Temporal Alignment Module: </p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">ii) For each patient 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi mathvariant="script">
              D 
            </mi> 
            <mi>
              i 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mi>
              Y 
            </mi> 
            <mi>
              i 
            </mi> 
           </msub> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           ∈ 
         </mo> 
         <mi>
           ℬ 
         </mi> 
        </mrow> 
       </math>: </p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">A) Encode visits: 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi>
            z 
          </mi> 
          <mi>
            n 
          </mi> 
         </msub> 
         <mo>
           = 
         </mo> 
         <msub> 
          <mi>
            E 
          </mi> 
          <mi>
            ϕ 
          </mi> 
         </msub> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msubsup> 
            <mi>
              X 
            </mi> 
            <mi>
              n 
            </mi> 
            <mrow> 
             <mi>
               M 
             </mi> 
             <mi>
               R 
             </mi> 
             <mi>
               I 
             </mi> 
            </mrow> 
           </msubsup> 
           <mo>
             , 
           </mo> 
           <msubsup> 
            <mi>
              X 
            </mi> 
            <mi>
              n 
            </mi> 
            <mrow> 
             <mi>
               P 
             </mi> 
             <mi>
               E 
             </mi> 
             <mi>
               T 
             </mi> 
            </mrow> 
           </msubsup> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math> for 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           n 
         </mi> 
         <mo>
           = 
         </mo> 
         <mn>
           1 
         </mn> 
         <mo>
           , 
         </mo> 
         <mn>
           ... 
         </mn> 
         <mo>
           , 
         </mo> 
         <msub> 
          <mi>
            N 
          </mi> 
          <mi>
            i 
          </mi> 
         </msub> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">B) Solve ODE: 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           z 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mi>
            t 
          </mi> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           = 
         </mo> 
         <mtext>
           ODESolve 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              f 
            </mi> 
            <mi>
              θ 
            </mi> 
           </msub> 
           <mo>
             , 
           </mo> 
           <msub> 
            <mi>
              z 
            </mi> 
            <mn>
              1 
            </mn> 
           </msub> 
           <mo>
             , 
           </mo> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                t 
              </mi> 
              <mn>
                1 
              </mn> 
             </msub> 
             <mo>
               , 
             </mo> 
             <mn>
               ... 
             </mn> 
             <mo>
               , 
             </mo> 
             <msub> 
              <mi>
                t 
              </mi> 
              <mrow> 
               <msub> 
                <mi>
                  N 
                </mi> 
                <mi>
                  i 
                </mi> 
               </msub> 
              </mrow> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">C) Interpolate: 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mstyle mathvariant="bold" mathsize="normal"> 
           <mi>
             Z 
           </mi> 
          </mstyle> 
          <mrow> 
           <mtext>
             aligned 
           </mtext> 
          </mrow> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mrow> 
          <mo>
            [ 
          </mo> 
          <mrow> 
           <mi>
             z 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                τ 
              </mi> 
              <mn>
                1 
              </mn> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mo>
             , 
           </mo> 
           <mn>
             ... 
           </mn> 
           <mo>
             , 
           </mo> 
           <mi>
             z 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                τ 
              </mi> 
              <mi>
                T 
              </mi> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ] 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">iii) Prediction &amp; Regularization:</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">iv) Encode: 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           H 
         </mi> 
         <mo>
           = 
         </mo> 
         <mtext>
           Transformer 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             TCN 
           </mtext> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msub> 
              <mi>
                Z 
              </mi> 
              <mrow> 
               <mtext>
                 aligned 
               </mtext> 
              </mrow> 
             </msub> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">v) Compute attention: 
       <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mi>
           A 
         </mi> 
         <mo>
           = 
         </mo> 
         <mi>
           σ 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              g 
            </mi> 
            <mi>
              ψ 
            </mi> 
           </msub> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mi>
              H 
            </mi> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">vi) Compute prediction: 
       <math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"> <mrow> 
         <mover accent="true"> 
          <mi>
            Y 
          </mi> 
          <mo>
            ^ 
          </mo> 
         </mover> 
         <mo>
           = 
         </mo> 
         <mtext>
           Classifier 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             GlobalAvgPool 
           </mtext> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mi>
               A 
             </mi> 
             <mo>
               ⊙ 
             </mo> 
             <mi>
               H 
             </mi> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">vii) Generate counterfactuals 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             Z 
           </mi> 
           <mo>
             ˜ 
           </mo> 
          </mover> 
          <mrow> 
           <mtext>
             aligned 
           </mtext> 
          </mrow> 
         </msub> 
        </mrow> 
       </math> and prediction 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mover accent="true"> 
         <mi>
           Y 
         </mi> 
         <mo>
           ˜ 
         </mo> 
        </mover> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">viii) Update Parameters:</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">ix) 
       <math xmlns="http://www.w3.org/1998/Math/MathML" display="inline"> <mrow> 
         <msub> 
          <mi>
            ℒ 
          </mi> 
          <mrow> 
           <mtext>
             total 
           </mtext> 
          </mrow> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mtext>
           BCE 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mover accent="true"> 
            <mi>
              Y 
            </mi> 
            <mo>
              ^ 
            </mo> 
           </mover> 
           <mo>
             , 
           </mo> 
           <mi>
             Y 
           </mi> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
         <mo>
           + 
         </mo> 
         <mi>
           β 
         </mi> 
         <mo>
           ⋅ 
         </mo> 
         <mtext>
           KL 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mi>
             P 
           </mi> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <mover accent="true"> 
              <mi>
                Y 
              </mi> 
              <mo>
                ^ 
              </mo> 
             </mover> 
             <mrow> 
              <mo>
                | 
              </mo> 
              <mi>
                Z 
              </mi> 
             </mrow> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
           <mrow> 
            <mo>
              ‖ 
            </mo> 
            <mrow> 
             <mi>
               P 
             </mi> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <mover accent="true"> 
                <mi>
                  Y 
                </mi> 
                <mo>
                  ˜ 
                </mo> 
               </mover> 
               <mrow> 
                <mo>
                  | 
                </mo> 
                <mover accent="true"> 
                 <mi>
                   Z 
                 </mi> 
                 <mo>
                   ˜ 
                 </mo> 
                </mover> 
               </mrow> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">x) 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <mtext>
           Θ 
         </mtext> 
         <mo>
           ← 
         </mo> 
         <mtext>
           Θ 
         </mtext> 
         <mo>
           − 
         </mo> 
         <mi>
           η 
         </mi> 
         <msub> 
          <mo>
            ∇ 
          </mo> 
          <mtext>
            Θ 
          </mtext> 
         </msub> 
         <msub> 
          <mi>
            ℒ 
          </mi> 
          <mrow> 
           <mtext>
             total 
           </mtext> 
          </mrow> 
         </msub> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="custom-bottom-td aleft" width="100.00%"><p style="text-align:left">3) Return trained parameters 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mi>
          Θ 
        </mi> 
       </math> </p></td> 
    </tr> 
    <tr> 
     <td class="custom-bottom-td custom-top-td aleft" width="100.00%"><p style="text-align:left">Algorithm 2 CASCADE-Net Inference Procedure</p></td> 
    </tr> 
    <tr> 
     <td class="custom-top-td aleft" width="100.00%"><p style="text-align:left">Input: New patient data 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi mathvariant="script">
            D 
          </mi> 
          <mtext>
            * 
          </mtext> 
         </msub> 
         <mo>
           = 
         </mo> 
         <msubsup> 
          <mrow> 
           <mrow> 
            <mo>
              { 
            </mo> 
            <mrow> 
             <mrow> 
              <mo>
                ( 
              </mo> 
              <mrow> 
               <msub> 
                <mi>
                  t 
                </mi> 
                <mi>
                  n 
                </mi> 
               </msub> 
               <mo>
                 , 
               </mo> 
               <msubsup> 
                <mi>
                  X 
                </mi> 
                <mi>
                  n 
                </mi> 
                <mrow> 
                 <mi>
                   M 
                 </mi> 
                 <mi>
                   R 
                 </mi> 
                 <mi>
                   I 
                 </mi> 
                </mrow> 
               </msubsup> 
               <mo>
                 , 
               </mo> 
               <msubsup> 
                <mi>
                  X 
                </mi> 
                <mi>
                  n 
                </mi> 
                <mrow> 
                 <mi>
                   P 
                 </mi> 
                 <mi>
                   E 
                 </mi> 
                 <mi>
                   T 
                 </mi> 
                </mrow> 
               </msubsup> 
              </mrow> 
              <mo>
                ) 
              </mo> 
             </mrow> 
            </mrow> 
            <mo>
              } 
            </mo> 
           </mrow> 
          </mrow> 
          <mrow> 
           <mi>
             n 
           </mi> 
           <mo>
             = 
           </mo> 
           <mn>
             1 
           </mn> 
          </mrow> 
          <mrow> 
           <msub> 
            <mi>
              N 
            </mi> 
            <mtext>
              * 
            </mtext> 
           </msub> 
          </mrow> 
         </msubsup> 
        </mrow> 
       </math>, Trained parameters 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mtext>
          Θ 
        </mtext> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">Output: Prediction probability 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             Y 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mtext>
            * 
          </mtext> 
         </msub> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">Summary: 1) Align the new patient’s timeline.</p><p style="text-align:left">2) Predict using the trained model.</p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">1) Encode and solve ODE for 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mi mathvariant="script">
            D 
          </mi> 
          <mtext>
            * 
          </mtext> 
         </msub> 
        </mrow> 
       </math> to get 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msubsup> 
          <mi>
            Z 
          </mi> 
          <mrow> 
           <mtext>
             aligned 
           </mtext> 
          </mrow> 
          <mtext>
            * 
          </mtext> 
         </msubsup> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">2) 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msup> 
          <mstyle mathvariant="bold" mathsize="normal"> 
           <mi>
             H 
           </mi> 
          </mstyle> 
          <mtext>
            * 
          </mtext> 
         </msup> 
         <mo>
           = 
         </mo> 
         <mtext>
           Transformer 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             TCN 
           </mtext> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msubsup> 
              <mstyle mathvariant="bold" mathsize="normal"> 
               <mi>
                 Z 
               </mi> 
              </mstyle> 
              <mrow> 
               <mtext>
                 aligned 
               </mtext> 
              </mrow> 
              <mtext>
                * 
              </mtext> 
             </msubsup> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">3) 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msup> 
          <mi>
            A 
          </mi> 
          <mtext>
            * 
          </mtext> 
         </msup> 
         <mo>
           = 
         </mo> 
         <mi>
           σ 
         </mi> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <msub> 
            <mi>
              g 
            </mi> 
            <mi>
              ψ 
            </mi> 
           </msub> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msup> 
              <mstyle mathvariant="bold" mathsize="normal"> 
               <mi>
                 H 
               </mi> 
              </mstyle> 
              <mtext>
                * 
              </mtext> 
             </msup> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="aleft" width="100.00%"><p style="text-align:left">4) 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             Y 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mtext>
            * 
          </mtext> 
         </msub> 
         <mo>
           = 
         </mo> 
         <mtext>
           Classifier 
         </mtext> 
         <mrow> 
          <mo>
            ( 
          </mo> 
          <mrow> 
           <mtext>
             GlobalAvgPool 
           </mtext> 
           <mrow> 
            <mo>
              ( 
            </mo> 
            <mrow> 
             <msup> 
              <mi>
                A 
              </mi> 
              <mtext>
                * 
              </mtext> 
             </msup> 
             <mo>
               ⊙ 
             </mo> 
             <msup> 
              <mstyle mathsize="normal" mathvariant="bold"> 
               <mi>
                 H 
               </mi> 
              </mstyle> 
              <mtext>
                * 
              </mtext> 
             </msup> 
            </mrow> 
            <mo>
              ) 
            </mo> 
           </mrow> 
          </mrow> 
          <mo>
            ) 
          </mo> 
         </mrow> 
        </mrow> 
       </math></p></td> 
    </tr> 
    <tr> 
     <td class="custom-bottom-td aleft" width="100.00%"><p style="text-align:left">5) Return 
       <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
         <msub> 
          <mover accent="true"> 
           <mi>
             Y 
           </mi> 
           <mo>
             ^ 
           </mo> 
          </mover> 
          <mtext>
            * 
          </mtext> 
         </msub> 
        </mrow> 
       </math></p></td> 
    </tr> 
   </table>
  </sec><sec id="s8">
   <title>
    <xref ref-type="bibr" rid="scirp.146372-"></xref>8. Discussion and Conclusion</title>
   <sec id="s8_1">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>8.1. Discussion</title>
    <p>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>CASCADE-Net provides a comprehensive solution to several key challenges in longitudinal medical image analysis. The Neural ODE-based alignment offers a principled, continuous-time approach to handling irregular sampling, superior to discrete RNN-based approximations. The integration of a counterfactual constraint within the learning objective is a significant step towards building more causally-aware AI models for healthcare. It moves beyond correlation towards learning invariant mechanisms, which should theoretically improve generalization across different hospitals and scanner protocols. The dynamic attention maps provided by CASCADE-Net offer tangible clinical utility beyond a simple prognostic score. For instance, a clinician reviewing a case with high predicted conversion risk ( 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <mover accent="true"> 
        <mi>
          Y 
        </mi> 
        <mo>
          ^ 
        </mo> 
       </mover> 
       <mo>
         &gt; 
       </mo> 
       <mn>
         0.8 
       </mn> 
      </mrow> 
     </math>) could examine the attention evolution over the 24-month timeline. If the maps show progressively increasing attention in the hippocampus and entorhinal cortex—regions known to be affected early in AD—this objective, data-driven visualization could support a decision to shorten the next monitoring interval from 12 to 6 months. Conversely, if the attention pattern is diffuse or stable, even with a moderate risk score, a clinician might maintain a standard monitoring schedule. This interpretable output transforms the model from a black-box predictor into a decision-support tool that provides a clear rationale for clinical actions. A limitation of our current work is the definition of the confounder distribution 
     <math xmlns="http://www.w3.org/1998/Math/MathML"> <mrow> 
       <msub> 
        <mi>
          p 
        </mi> 
        <mi>
          c 
        </mi> 
       </msub> 
      </mrow> 
     </math>. In this study, we modeled it as a simple Gaussian based on population statistics for age-related change. Furthermore, our model was developed and validated exclusively on data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) cohort. While ADNI provides high-quality, well-curated data, its specific inclusion criteria and demographic profile may limit the immediate generalizability of our findings to more diverse, real-world clinical populations with different ethnic backgrounds, comorbidities, and imaging protocols. Future work will focus on learning more sophisticated confounder models from large-scale control data and, crucially, validating the CASCADE-Net framework on independent, multi-site datasets to confirm its robustness and clinical utility across diverse healthcare settings.</p>
   </sec>
   <sec id="s8_2">
    <title>
     <xref ref-type="bibr" rid="scirp.146372-"></xref>8.2. Conclusion</title>
    <p>We presented CASCADE-Net, a novel end-to-end pipeline for predicting MCI-to-AD conversion. By integrating Neural ODEs for dynamic temporal alignment with a counterfactually-regularized attention mechanism, our model achieves state-of-the-art prognostic performance. More importantly, it provides interpretable, spatio-temporal explanations that are grounded in causal reasoning, making its predictions more trustworthy and actionable for clinical practice. This architecture is general and can be adapted to other neurodegenerative diseases and longitudinal analysis tasks.</p>
   </sec>
  </sec><sec id="s9">
   <title>Acknowledgements</title>
   <p>Data used in preparation of this article were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu). The authors acknowledge that no funding has been recieved for this study. Its inspired by the need to help solve problems of the older people.</p>
  </sec>
 </body><back>
  <ref-list>
   <title>References</title>
   <ref id="scirp.146372-ref1">
    <label>1</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Petersen, R.C., Smith, G.E., Waring, S.C., Ivnik, R.J., Tangalos, E.G. and Kokmen, E. (1999) Mild Cognitive Impairment: Clinical Characterization and Outcome. Archives of Neurology, 56, 303-308. &gt;https://doi.org/10.1001/archneur.56.3.303
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref2">
    <label>2</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jack, C.R., Wiste, H.J., Vemuri, P., Weigand, S.D., Senjem, M.L., Zeng, G., et al. (2010) Brain Beta-Amyloid Measures and Magnetic Resonance Imaging Atrophy Both Predict Time-to-Progression from Mild Cognitive Impairment to Alzheimer’s Disease. Brain, 133, 3336-3348. &gt;https://doi.org/10.1093/brain/awq277
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref3">
    <label>3</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Li, X., Wang, X., Su, L., Hu, X. and Zhou, Y. (2018) Prediction of Conversion from Mild Cognitive Impairment to Alzheimer’s Disease Dementia Based upon Biomarkers and Neuropsychological Test Performance. Neurobiology of Aging, 66, 120-128.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref4">
    <label>4</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Geirhos, R., Jacobsen, J., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., et al. (2020) Shortcut Learning in Deep Neural Networks. Nature Machine Intelligence, 2, 665-673. &gt;https://doi.org/10.1038/s42256-020-00257-z
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref5">
    <label>5</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Cherubini, A., Peran, P., Spoletini, I., Di Paola, M., Di Iulio, F., Hagberg, G.E., et al. (2010) Combined MRI-Based Hippocampal Volumetry and 1h-mrs in Mild Cognitive Impairment: A Preliminary Study. Neuroradiology, 52, 503-511.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref6">
    <label>6</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Liu, X., Shen, L., Liu, J., Zhang, J. and Li, G. (2015) A Deep Convolutional Neural Network-Based Regression Method for 3D Patchwise Hippocampus Segmentation. Informatics in Medicine Unlocked, 1, 1-7.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref7">
    <label>7</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Nguyen, M., He, T., An, L., Alexander, D.C., Feng, J. and Yeo, B.T.T. (2017) Long-Term Memory Modeling with Neural Networks for Predicting Alzheimer’s Disease Progression. In: International Workshop on Machine Learning in Medical Imaging, Springer, 250-258.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref8">
    <label>8</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Neil, D., Pfeiffer, M. and Liu, S.-C. (2016) Phased LSTM: Accelerating Recurrent Network Training for Long or Event-Based Sequences. 2016 Conference on Neural Information Processing Systems, Barcelona, 5-10 December 2016, 3882-3890.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref9">
    <label>9</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Che, Z., Purushotham, S., Cho, K., Sontag, D. and Liu, Y. (2018) Recurrent Neural Networks for Multivariate Time Series with Missing Values. Scientific Reports, 8, Article No. 6085. &gt;https://doi.org/10.1038/s41598-018-24271-9
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref10">
    <label>10</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Chen, R.T., Rubanova, Y., Bettencourt, J. and Duvenaud, D.K. (2018) Neural Ordinary Differential Equations. 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, 3-8 December 2018. &gt;https://proceedings.neurips.cc/paper/2018/hash/69386f6bb1dfed68692a24c8686939b9-Abstract.html 
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref11">
    <label>11</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Jia, J. and Benson, A.R. (2019) Neural Odes for Informative Missingness in Multivariate Time Series.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref12">
    <label>12</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł. and Polosukhin, I. (2017) Attention Is All You Need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, 4-9 December 2017, 6000-6010.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref13">
    <label>13</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Korolev, S., Safiullin, A., Belyaev, M. and Dodonova, Y. (2020) Residual and Plain Convolutional Neural Networks for 3d Brain MRI Classification. IEEE Journal of Biomedical and Health Informatics, 25, 743-752.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref14">
    <label>14</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Pearl, J. (2009) Causality. Cambridge University Press. &gt;https://doi.org/10.1017/cbo9780511803161
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref15">
    <label>15</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Arjovsky, M., Bottou, L., Gulrajani, I. and Lopez-Paz, D. (2019) Invariant Risk Minimization.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref16">
    <label>16</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Schrouff, J., Monteiro, J.M., Ferreira, C., Rosa, M.J., Wardle, J., Whyte, C., et al. (2019) Learning the Super-Resolution Imaging of Structural Magnetic Resonances with Counterfactual Modeling.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref17">
    <label>17</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Bai, S., Kolter, J.Z. and Koltun, V. (2018) An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref18">
    <label>18</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Weiner, M.W., Veitch, D.P., Aisen, P.S., Beckett, L.A., Cairns, N.J., Green, R.C., et al. (2013) The Alzheimer’s Disease Neuroimaging Initiative: A Review of Papers Published since Its Inception. Alzheimer’s&amp;Dementia, 9, e111-e194.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref19">
    <label>19</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., et al. (2019) PyTorch: An Imperative Style, High-Performance Deep Learning Library. Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, 8-14 December 2019, 8026-8037.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref20">
    <label>20</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Kingma, D.P. and Ba, J. (2014) Adam: A Method for Stochastic Optimization.
    </mixed-citation>
   </ref>
   <ref id="scirp.146372-ref21">
    <label>21</label>
    <mixed-citation publication-type="other" xlink:type="simple">
     Braak, H. and Braak, E. (1991) Neuropathological Stageing of Alzheimer-Related Changes. Acta Neuropathologica, 82, 239-259. &gt;https://doi.org/10.1007/bf00308809
    </mixed-citation>
   </ref>
  </ref-list>
 </back>
</article>