TITLE:
Development and Temporal Validation of Machine Learning Models for Predicting High Cardiac Biomarker Burden in Patients Undergoing Hemodialysis
AUTHORS:
Tong Wang, Jingchun Yang, Xiaoxiao Zhang, Jing Wang, Lin Li
KEYWORDS:
Maintenance Hemodialysis, Cardiac Biomarkers, Machine Learning, Stacking Ensemble, Temporal Validation, Contemporaneous Classification, Explainable Artificial Intelligence
JOURNAL NAME:
Journal of Biosciences and Medicines,
Vol.14 No.9,
September
10,
2026
ABSTRACT: Background: Cardiovascular assessment in patients undergoing maintenance hemodialysis is complicated by elevations and variability in cardiac biomarkers. We aimed to develop and temporally validate an interpretable machine learning model for classifying a study-defined composite of elevated B-type natriuretic peptide (BNP) or high-sensitivity cardiac troponin I (hs-cTnI) at the same clinical assessment episode. Methods: This study included 960 adults receiving maintenance hemodialysis in China. Patients treated from January 2021 to December 2023 comprised the development-period cohort, and those treated from January 2024 to December 2025 comprised the temporal validation cohort. The development-period cohort was divided into training and same-period holdout sets. Least absolute shrinkage and selection operator (LASSO) regression selected predictors, and six models were developed. Three top-performing models were combined in a stacking ensemble. Discrimination, grouped calibration, Brier score, decision-curve analysis, and Shapley additive explanations (SHAP) were evaluated. Results: LASSO retained 13 predictors. The stacking ensemble achieved ROC-AUC point estimates of 0.848, 0.863, and 0.835 in training out-of-fold evaluation, same-period holdout evaluation, and temporal validation, respectively. Corresponding precision-recall AUC point estimates were 0.77, 0.76, and 0.70, and Brier scores were 0.15, 0.15, and 0.16. In temporal validation, sensitivity, specificity, and accuracy point estimates were 0.75, 0.76, and 0.75. SHAP identified prognostic nutritional index (PNI), log-transformed D-dimer, age, red cell distribution width (RDW), and beta-2 microglobulin as leading predictors. Small between-model differences should be interpreted cautiously because confidence intervals were not estimated. Conclusions: The stacking ensemble showed comparable point-estimate performance across same-period holdout and temporal validation data and may support further investigation of contemporaneous biomarker-burden classification. Multicenter external validation, standardized sampling, uncertainty quantification, and prospective evaluation are required before clinical implementation.