Diabetes mellitus, particularly type 2 diabetes, poses a growing global health challenge, with rising prevalence in low- and middle-income countries where diagnostic services are limited. Early detection is critical, but traditional screening methods often miss at-risk individuals.
A recent study published in Scientific Reports (2025) evaluated machine learning models to classify diabetes using sociodemographic, behavioral, and clinical predictors. The researchers analyzed data from the National Health and Nutrition Examination Survey (NHANES) spanning 2013–2018, including 9,096 participants.
The study compared several algorithms, including logistic regression, random forest, and XGBoost. The random forest model achieved the highest accuracy (0.80) and area under the ROC curve (0.87), with key predictors being age, waist circumference, and triglyceride levels.
These findings suggest that machine learning can effectively identify individuals at risk for diabetes using readily available data, potentially improving screening in resource-limited settings. The authors emphasize the need for external validation before clinical implementation.