Background: Diabetes mellitus is a major public health concern in India, with increasing prevalence driven by demographic, metabolic, behavioral, and lifestyle-related factors. Early identification of individuals at high risk of diabetes can support timely prevention and improve health outcomes. Machine learning (ML) approaches offer opportunities to identify complex and nonlinear relationships among multiple risk factors and may complement conventional statistical approaches. Objective: This study aimed to investigate and compare the performance of different ML models for predicting diabetes risk, with particular emphasis on identifying the most influential predictors and assessing the potential applicability of ML-based approaches for early risk stratification in Indian populations. Methods: A cross-sectional analytical approach was used to evaluate demographic, physiological, and lifestyle-related variables associated with diabetes risk. The study considered variables including age, body mass index (BMI), blood glucose, blood pressure, insulin, family history of diabetes, and physical activity. Data preprocessing included missing-value management, outlier identification, categorical encoding, and feature standardization. The dataset was divided into training (70%) and testing (30%) subsets using stratified sampling. Seven supervised ML algorithms—Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, Extreme Gradient Boosting (XGBoost), and Artificial Neural Network—were compared. Model performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC-ROC). Pearson correlation, one-way ANOVA, and multivariate logistic regression were additionally used to examine associations between predictors and diabetes risk. Results: XGBoost demonstrated the strongest overall predictive performance, achieving an accuracy of 89.2%, precision of 0.88, recall of 0.87, F1-score of 0.87, and AUC of 0.93. Random Forest and the neural network also demonstrated strong performance, with accuracy of 87.6% and 88.5% and AUCs of 0.91 and 0.92, respectively. Glucose level was the most influential predictor, followed by BMI and age. Statistical analyses further supported the importance of these metabolic and demographic factors in diabetes risk prediction. Conclusion: Ensemble and advanced ML approaches, particularly XGBoost and Random Forest, demonstrated promising performance for diabetes risk prediction. Integrating ML with statistical analysis and interpretable AI approaches may strengthen early risk identification and support evidence-based diabetes prevention and personalized healthcare strategies in Indian populations. Further validation using larger, diverse, and multi center datasets is required before clinical implementation.
| Published in | Machine Learning Research (Volume 11, Issue 2) |
| DOI | 10.11648/j.mlr.20261102.12 |
| Page(s) | 76-84 |
| Creative Commons |
This is an Open Access article, distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution and reproduction in any medium or format, provided the original work is properly cited. |
| Copyright |
Copyright © The Author(s), 2026. Published by Science Publishing Group |
Diabetes Early Prediction, Machine Learning, India, Random Forest, Xgboost, Nfhs-5
Variable | Type | Description |
|---|---|---|
Age | Numeric | Age in years |
BMI | Numeric | Body Mass Index |
Glucose | Numeric | Blood glucose level |
Blood Pressure | Numeric | Diastolic blood pressure |
Insulin | Numeric | Serum insulin |
Family History | Categorical | Diabetes in family (Yes/No) |
Physical Activity | Categorical | Active/Sedentary |
Outcome | Binary | Diabetic (1) / non-diabetic (0) |
Model | Accuracy (%) | Precision | Recall | F1-score | AUC |
|---|---|---|---|---|---|
Logistic Regression | 78.5 | 0.76 | 0.74 | 0.75 | 0.81 |
Decision Tree | 80.2 | 0.78 | 0.77 | 0.77 | 0.83 |
SVM | 82.4 | 0.81 | 0.79 | 0.8 | 0.85 |
KNN | 79.8 | 0.77 | 0.76 | 0.76 | 0.82 |
Random Forest | 87.6 | 0.86 | 0.85 | 0.85 | 0.91 |
XGBoost | 89.2 | 0.88 | 0.87 | 0.87 | 0.93 |
Neural Network | 88.5 | 0.87 | 0.86 | 0.86 | 0.92 |
Rank | Feature | Importance Score |
|---|---|---|
1 | Glucose Level | 0.32 |
2 | BMI | 0.21 |
3 | Age | 0.18 |
4 | Blood Pressure | 0.12 |
5 | Family History | 0.09 |
6 | Physical Activity | 0.08 |
AI | Artificial Intelligence |
ANN | Artificial Neural Network |
ANOVA | Analysis of Variance |
AUC | Area Under the Curve |
AUC-ROC | Area Under the Receiver Operating Characteristic Curve |
BMI | Body Mass Index |
CNN | Convolutional Neural Network |
DL | Deep Learning |
DT | Decision Tree |
EHRs | Electronic Health Records |
F1-score | F1 Score |
IQR | Interquartile Range |
KNN | K-Nearest Neighbors |
LR | Logistic Regression |
ML | Machine Learning |
NFHS-5 | National Family Health Survey – 5 |
OR | Odds Ratio |
RF | Random Forest |
RFE | Recursive Feature Elimination |
ROC | Receiver Operating Characteristic |
RNN | Recurrent Neural Network |
SVM | Support Vector Machine |
SPSS | Statistical Package for the Social Sciences |
XAI | Explainable Artificial Intelligence |
XGBoost | Extreme Gradient Boosting |
| [1] | Banday, M. Z., Sameer, A. S., & Nissar, S. (2020). Pathophysiology of diabetes: An overview. Avicenna Journal of Medicine, 10(4), 174–188. |
| [2] | Pradeepa, R., & Mohan, V. (2021). Epidemiology of type 2 diabetes in India. Indian Journal of Ophthalmology, 69(11), 2932–2938. |
| [3] | Deberneh, H. M., & Kim, I. (2021). Prediction of type 2 diabetes based on machine learning algorithm. International Journal of Environmental Research and Public Health, 18(6), 3317. |
| [4] | Fregoso-Aparicio, L., Noguez, J., Montesinos, L., & García-García, J. A. (2021). Machine learning and deep learning predictive models for type 2 diabetes: A systematic review. Diabetology & Metabolic Syndrome, 13, 148. |
| [5] | Chang, V., Bailey, J., Xu, Q. A., & Sun, Z. (2022). Pima Indians diabetes mellitus classification based on machine learning (ML) algorithms. Neural Computing and Applications, 35, 16157–16173. |
| [6] | Naz, H., & Ahuja, S. (2020). Deep learning approach for diabetes prediction using PIMA Indian dataset. Journal of Diabetes & Metabolic Disorders, 19(1), 391–403. |
| [7] | Ali, S., Abuhmed, T., El-Sappagh, S., Muhammad, K., Alonso-Moral, J. M., Confalonieri, R., Guidotti, R., Del Ser, J., Díaz-Rodríguez, N., & Herrera, F. (2023). Explainable artificial intelligence (XAI): What we know and what is left to attain trustworthy artificial intelligence. Information Fusion, 99, 101805. |
| [8] | Wang, L., Wang, X., Chen, A., Jin, X., & Che, H. (2020). Prediction of type 2 diabetes risk and its effect evaluation based on the XGBoost model. Healthcare, 8(3), 247. |
| [9] | Liu, K., Li, L., Ma, Y., Jiang, J., Liu, Z., Ye, Z., Liu, S., Pu, C., Chen, C., & Wan, Y. (2023). Machine learning models for blood glucose level prediction in patients with diabetes mellitus: Systematic review and network meta-analysis. JMIR Medical Informatics, 11, e47833. |
| [10] | Tan, K. R., Seng, J. J. B., Kwan, Y. H., Chen, Y. J., Zainudin, S. B., Loh, D. H. F., Liu, N., & Low, L. L. (2023). Evaluation of machine learning methods developed for prediction of diabetes complications: A systematic review. Journal of Diabetes Science and Technology, 17(2), 474–489. |
| [11] | Ansari, G. A., Shafi, S., Ansari, M. D., & Shadab, A. (2025). Advanced supervised machine learning methods for precise diabetes mellitus prediction using feature selection. Frontiers in Medicine, 12, 1620268. |
| [12] | Natarajan, K., Baskaran, D., & Kamalanathan, S. (2025). An adaptive ensemble feature selection technique for model-agnostic diabetes prediction. Scientific Reports, 15, 6907. |
| [13] | Prakash, E. P., Srihari, K., Karthik, S., Kamal, M. V., Dileep, P., Reddy, B. S., Mukunthan, M. A., Somasundaram, K., Jaikumar, R., Gayathri, N., & Sahile, K. (2022). Implementation of artificial neural network to predict diabetes with high-quality health system. Computational and Mathematical Methods in Medicine, 2022, 1174173. |
| [14] | Yu, W., Liu, T., Valdez, R., Gwinn, M., & Khoury, M. J. (2010). Application of support vector machine modeling for prediction of common diseases: The case of diabetes and pre-diabetes. BMC Medical Informatics and Decision Making, 10, 16. |
| [15] | Rahman, F., Hossain, S., Tiang, J.-J., & Nahid, A.-A. (2025). Diabetes prediction using feature selection algorithms and boosting-based machine learning classifiers. Diagnostics, 15(20), 2622. |
| [16] | Tasin, I., Nabil, T. U., Islam, S., & Khan, R. (2022). Diabetes prediction using machine learning and explainable AI techniques. Healthcare Technology Letters, 10(1–2), 1–10. |
| [17] | Loh, H. W., Ooi, C. P., Seoni, S., Barua, P. D., Molinari, F., & Acharya, U. R. (2022). Application of explainable artificial intelligence for healthcare: A systematic review of the last decade (2011–2022). Computer Methods and Programs in Biomedicine, 226, 107161. |
| [18] | Nicolucci, A., Romeo, L., Bernardini, M., Vespasiani, M., Rossi, M. C., Petrelli, M., Ceriello, A., Di Bartolo, P., Frontoni, E., & Vespasiani, G. (2022). Prediction of complications of type 2 diabetes: A machine learning approach. Diabetes Research and Clinical Practice, 190, 110013. |
| [19] | Breiman, L. (2001). Random forests. Machine Learning, 45, 5–32. |
| [20] | Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20, 273–297. |
| [21] | Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). Association for Computing Machinery. |
| [22] | Zou, Q., Qu, K., Luo, Y., Yin, D., Ju, Y., & Tang, H. (2018). Predicting diabetes mellitus with machine learning techniques. Frontiers in Genetics, 9, 515. |
APA Style
Singh, G. (2026). Investigation on Machine Learning Models for Predicting Diabetes Risk in Indian Populations. Machine Learning Research, 11(2), 76-84. https://doi.org/10.11648/j.mlr.20261102.12
ACS Style
Singh, G. Investigation on Machine Learning Models for Predicting Diabetes Risk in Indian Populations. Mach. Learn. Res. 2026, 11(2), 76-84. doi: 10.11648/j.mlr.20261102.12
AMA Style
Singh G. Investigation on Machine Learning Models for Predicting Diabetes Risk in Indian Populations. Mach Learn Res. 2026;11(2):76-84. doi: 10.11648/j.mlr.20261102.12
@article{10.11648/j.mlr.20261102.12,
author = {Gajendra Singh},
title = {Investigation on Machine Learning Models for Predicting Diabetes Risk in Indian Populations},
journal = {Machine Learning Research},
volume = {11},
number = {2},
pages = {76-84},
doi = {10.11648/j.mlr.20261102.12},
url = {https://doi.org/10.11648/j.mlr.20261102.12},
eprint = {https://article.sciencepublishinggroup.com/pdf/10.11648.j.mlr.20261102.12},
abstract = {Background: Diabetes mellitus is a major public health concern in India, with increasing prevalence driven by demographic, metabolic, behavioral, and lifestyle-related factors. Early identification of individuals at high risk of diabetes can support timely prevention and improve health outcomes. Machine learning (ML) approaches offer opportunities to identify complex and nonlinear relationships among multiple risk factors and may complement conventional statistical approaches. Objective: This study aimed to investigate and compare the performance of different ML models for predicting diabetes risk, with particular emphasis on identifying the most influential predictors and assessing the potential applicability of ML-based approaches for early risk stratification in Indian populations. Methods: A cross-sectional analytical approach was used to evaluate demographic, physiological, and lifestyle-related variables associated with diabetes risk. The study considered variables including age, body mass index (BMI), blood glucose, blood pressure, insulin, family history of diabetes, and physical activity. Data preprocessing included missing-value management, outlier identification, categorical encoding, and feature standardization. The dataset was divided into training (70%) and testing (30%) subsets using stratified sampling. Seven supervised ML algorithms—Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, Extreme Gradient Boosting (XGBoost), and Artificial Neural Network—were compared. Model performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC-ROC). Pearson correlation, one-way ANOVA, and multivariate logistic regression were additionally used to examine associations between predictors and diabetes risk. Results: XGBoost demonstrated the strongest overall predictive performance, achieving an accuracy of 89.2%, precision of 0.88, recall of 0.87, F1-score of 0.87, and AUC of 0.93. Random Forest and the neural network also demonstrated strong performance, with accuracy of 87.6% and 88.5% and AUCs of 0.91 and 0.92, respectively. Glucose level was the most influential predictor, followed by BMI and age. Statistical analyses further supported the importance of these metabolic and demographic factors in diabetes risk prediction. Conclusion: Ensemble and advanced ML approaches, particularly XGBoost and Random Forest, demonstrated promising performance for diabetes risk prediction. Integrating ML with statistical analysis and interpretable AI approaches may strengthen early risk identification and support evidence-based diabetes prevention and personalized healthcare strategies in Indian populations. Further validation using larger, diverse, and multi center datasets is required before clinical implementation.},
year = {2026}
}
TY - JOUR T1 - Investigation on Machine Learning Models for Predicting Diabetes Risk in Indian Populations AU - Gajendra Singh Y1 - 2026/09/02 PY - 2026 N1 - https://doi.org/10.11648/j.mlr.20261102.12 DO - 10.11648/j.mlr.20261102.12 T2 - Machine Learning Research JF - Machine Learning Research JO - Machine Learning Research SP - 76 EP - 84 PB - Science Publishing Group SN - 2637-5680 UR - https://doi.org/10.11648/j.mlr.20261102.12 AB - Background: Diabetes mellitus is a major public health concern in India, with increasing prevalence driven by demographic, metabolic, behavioral, and lifestyle-related factors. Early identification of individuals at high risk of diabetes can support timely prevention and improve health outcomes. Machine learning (ML) approaches offer opportunities to identify complex and nonlinear relationships among multiple risk factors and may complement conventional statistical approaches. Objective: This study aimed to investigate and compare the performance of different ML models for predicting diabetes risk, with particular emphasis on identifying the most influential predictors and assessing the potential applicability of ML-based approaches for early risk stratification in Indian populations. Methods: A cross-sectional analytical approach was used to evaluate demographic, physiological, and lifestyle-related variables associated with diabetes risk. The study considered variables including age, body mass index (BMI), blood glucose, blood pressure, insulin, family history of diabetes, and physical activity. Data preprocessing included missing-value management, outlier identification, categorical encoding, and feature standardization. The dataset was divided into training (70%) and testing (30%) subsets using stratified sampling. Seven supervised ML algorithms—Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, Extreme Gradient Boosting (XGBoost), and Artificial Neural Network—were compared. Model performance was assessed using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve (AUC-ROC). Pearson correlation, one-way ANOVA, and multivariate logistic regression were additionally used to examine associations between predictors and diabetes risk. Results: XGBoost demonstrated the strongest overall predictive performance, achieving an accuracy of 89.2%, precision of 0.88, recall of 0.87, F1-score of 0.87, and AUC of 0.93. Random Forest and the neural network also demonstrated strong performance, with accuracy of 87.6% and 88.5% and AUCs of 0.91 and 0.92, respectively. Glucose level was the most influential predictor, followed by BMI and age. Statistical analyses further supported the importance of these metabolic and demographic factors in diabetes risk prediction. Conclusion: Ensemble and advanced ML approaches, particularly XGBoost and Random Forest, demonstrated promising performance for diabetes risk prediction. Integrating ML with statistical analysis and interpretable AI approaches may strengthen early risk identification and support evidence-based diabetes prevention and personalized healthcare strategies in Indian populations. Further validation using larger, diverse, and multi center datasets is required before clinical implementation. VL - 11 IS - 2 ER -