Peter M. Socha, M. Oskoui, Jennifer A. Hutcheon, Sam Harper
2026.1.1ANNALS OF EPIDEMIOLOGY
tlooto Summary
A logistic regression model improved identification of cerebral palsy cases in administrative data, but residual misclassification remained.
Abstract
PURPOSE To improve the identification of cerebral palsy cases in administrative health data.
METHODS We included all children in a population-based cerebral palsy registry in Quebec, Canada, born from 1999-2002, and a sample of children without cerebral palsy. Population-based hospitalization and physician billing records through 2012 were obtained for all children. We used logistic regression to model the probability of cerebral palsy, using International Classification of Diseases codes for related diseases. We reported receiver operating characteristic (ROC) and precision-recall (PR) curves, and compared the accuracy to that of existing algorithms. We also reported the accuracy of cerebral palsy codes by age, data source, and gestational age at birth.
RESULTS The area under the ROC and PR curves of our model were 0.98 (95% CI: 0.97 to 0.99) and 0.73 (95% CI: 0.63 to 0.79), respectively. Cut-offs with a similar specificity to existing algorithms yielded sensitivities that were 1-14 percentage-points higher. The sensitivity of cerebral palsy codes was higher (and the specificity was lower) with longer follow-up times since birth, when using both hospitalization and billing records, and among children born preterm.
CONCLUSIONS Our model improved identification of cerebral palsy cases in administrative data, but residual misclassification remained.
Citation format
SOCHA, Peter M., et al. A multivariable model for improving the identification of cerebral palsy cases in administrative health data. ANNALS OF EPIDEMIOLOGY, 2026, 114: 26–31.