Open Access

G. Oliveira, Anderson O. Fonsêca, P. Rodrigues

2022.12.31Brazilian Journal of Biometrics

DOI: 10.28951/bjb.v40i4.605

tlooto Summary

This work aims to build a model as an alternative to traditional exams to identify the disease, using logistic regression, K-nearestneighbors, decision trees, random forest, and support vector machines for diabetes classification.

Abstract

Diabetes mellitus is one of the deadliest incurable diseases globally, and its cases continue upward. The identification of the disease in an early way helps fight it; however, blood tests can be considered invasive, discouraging its accomplishment. In this vein, this work aims to build a model as an alternative to traditional exams to identify the disease. Statistical learning algorithms such as logistic regression, K-nearestneighbors, decision trees, random forest, and support vector machines were used for diabetes classification. These models were considered separately and combined via hard and soft voting classifiers. Themethods were applied to a widely known dataset of 768 individuals and nine variables, compared using several accuracy metrics based on the confusion matrix, and used to estimate the probability of diabetesfor a given profile.

Citation format

OLIVEIRA, G.; FONSÊCA, Anderson O.; RODRIGUES, P. Diabetes diagnosis based on hard and soft voting classifiers combining statistical learning models. Brazilian Journal of Biometrics, 2022.