Machine Learning and Data ClassificationSimulation Techniques and ApplicationsStatistical Methods and Inference

Sun Y. Jeon, Sei J. Lee, Bocheng Jing, Alexandra K Lee, W. J. Deardorff, W. J. Boscardin

2026.2.28Communications in Statistics Case Studies Data Analysis and Applications

DOI: 10.1080/23737484.2026.2622374

Abstract

LASSO and backward selection methods were used to develop prediction models in electronic health record (EHR) data, varying the sample size (n) while holding the number of predictors (p) fixed. Utilizing Veterans Affairs (VA) EHR data from 2005 with vital status follow-up until 2017, we employed four modeling strategies across six sample sizes: LASSO with cross-validated λ and Bayesian information criterion (BIC) λ, backward selection using BIC, and p<.05 as the criterion. Six random samples (n= 1,264; 3,791; 12,636; 37,908; 63,180; 126,360) were generated from the full VA EHR data (∼1.26 million patient records) to predict 12-year mortality using 924 predictors. Discrimination and calibration of the developed models were evaluated in separate validation samples. For the two smallest samples (n= 1,264 and 3,791; n/p∼1 and 4), BIC LASSO yielded the most parsimonious and least overfit model. For larger samples, both LASSO and backward selection models demonstrated minimal overfitting and no evidence of miscalibration. However, BIC backward selection used far fewer predictors than the other methods. Our findings suggest that in small samples, BIC LASSO should be preferred as other methods yielded overfit models. In larger samples, both LASSO and backward selection appear to be reasonable methods for EHR prediction model development.

Citation format

JEON, Sun Y., et al. Comparing the accuracy and parsimony of LASSO and backward selection prediction models across varying sample sizes. Communications in Statistics Case Studies Data Analysis and Applications, 2026, 12(2): 157–168.