Machine Learning in HealthcareArtificial Intelligence in HealthcareECG Monitoring and Analysis

Prommin Buaphan, Nitipon Pongphaw

2026.5.20Journal of Bio-X Research

DOI: 10.34133/jbioxresearch.0091

Abstract

This study presents a systematic evaluation of classical machine learning, deep learning, and hybrid models for heart disease prediction using the University of California, Irvine dataset. A standardized framework with consistent preprocessing and 5-fold cross-validation was applied to ensure fair comparison. The results indicate that the hybrid models achieved competitive performance across multiple metrics, with the proposed convolutional neural network (CNN)–Attention–eXtreme Gradient Boosting (XGBoost) ensemble obtaining an accuracy of 82.07% ± 2.68% and an F1 score of 81.78% ± 2.68%. However, CNN + Support Vector Machine (SVM) achieved the highest accuracy (83.37% ± 2.11%) among all models, whereas Random Forest achieved the highest area under the curve (89.36% ± 1.42%) and showed strong calibration performance. In terms of probabilistic calibration, Brier Skill Score (BSS) analysis showed that all models outperformed the baseline (BSS > 0), with CNN + SVM (0.477), SVM (0.474), and Random Forest (0.465) achieving the highest calibration skill. Statistical analysis using the Friedman test and post hoc Wilcoxon signed-rank tests with Bonferroni correction revealed no statistically significant pairwise differences among the models, suggesting that the observed performance differences were marginal despite numerical variations. Computational analysis highlighted a trade-off between performance and efficiency, as hybrid models required longer training times (up to 7.59 s per fold) than lightweight tree-based models (<0.3 s). A cross-country evaluation further revealed reduced generalizability under domain shift.

Citation format

BUAPHAN, Prommin; PONGPHAW, Nitipon. Performance–efficiency trade-off analysis of hybrid deep learning models for heart disease prediction. Journal of Bio-X Research, 2026, 9.