An AI-Driven Framework for Diabetes Prediction and Classification Using Medical Data, SMOTE, and Grid Search CV Optimization

Authors

  • Devashish Chauhan Department Of Computer Science and Engineering, Graphic Era, Dehradun, India
  • Rakshit Sharma Department Of Computer Science and Engineering, Graphic Era, Dehradun, India
  • Shyam Ji Department Of Electronics and Communication, NIT Hamirpur, India
  • Suraj Singh Patwal Department Of Mechanical Engineering, IIT Patna, India

DOI:

https://doi.org/10.70112/ajeat-2026.15.2.4370

Keywords:

Supervised Machine Learning, Classification Task, Algorithms, SMOTE, Grid Search CV, Diabetes Prediction

Abstract

In India, approximately 101 million people are diabetic, while around 136 million are prediabetic, making early diagnosis essential for preventing severe health complications. The ML field of AI can emerge as an alternative for diagnosing diabetes. In this study, various algorithms are employed for diabetes prediction using the Pima Indian Diabetes Dataset. Hyperparameter optimization was performed using Grid Search Cross Validation (Grid Search CV) to identify the optimal model parameters. Since the dataset is imbalanced, experiments were conducted on both the original imbalanced dataset and a balanced synthesized dataset generated using SMOTE, which significantly reduced False Negative (FN) predictions across most algorithms. Furthermore, Random Forest and Naive Bayes achieved the highest recall of 72.73% after applying SMOTE, demonstrating improved effectiveness in identifying diabetic patients.

References

[1] A. Kumar Dewangan and P. Agrawal, “Classification of diabetes mellitus using machine learning techniques,” International Journal of Engineering and Applied Sciences, vol. 2, no. 5, p. 257905, 2015.

[2] A. Mujumdar and V. Vaidehi, “Diabetes prediction using machine learning algorithms,” Procedia Computer Science, vol. 165, pp. 292–299, 2019.

[3] C. S. M. Ali and O. S. Kareem, “Diabetes prediction using machine learning,” Asian Journal of Research in Computer Science, vol. 18, no. 6, pp. 89–109, 2025.

[4] P. B. Khokhar, V. Pentangelo, F. Palomba, and C. Gravino, “Towards transparent and accurate diabetes prediction using machine learning and explainable artificial intelligence,” arXiv preprint arXiv:2501.18071, 2025.

[5] S. Mondal and J. P. Choudhury, “Prediction of diabetes using machine learning models,” Science and Culture, 2024.

[6] N. P. Tigga and S. Garg, “Prediction of type 2 diabetes using machine learning classification methods,” Procedia Computer Science, vol. 167, pp. 706–716, 2020.

[7] M. K. Hasan, M. A. Alam, D. Das, E. Hossain, and M. Hasan, “Diabetes prediction using ensembling of different machine learning classifiers,” IEEE Access, vol. 8, pp. 76516–76531, 2020.

[8] M. E. Febrian, F. X. Ferdinan, G. P. Sendani, K. M. Suryanigrum, and R. Yunanda, “Diabetes prediction using supervised machine learning,” Procedia Computer Science, vol. 216, pp. 21–30, 2023.

[9] R. Yacouby and D. Axman, “Probabilistic extension of precision, recall, and F1 score for more thorough evaluation of classification models,” in Proc. 1st Workshop on Evaluation and Comparison of NLP Systems, 2020, pp. 79–91.

[10] K. M. Sujon, R. Hassan, K. Choi, and M. A. Samad, “Accuracy, precision, recall, F1-score, or MCC? Empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models,” Journal of Big Data, vol. 12, no. 1, p. 268, 2025.

[11] E. Helmud, F. Fitriyani, and P. Romadiana, “Classification comparison performance of supervised machine learning random forest and decision tree algorithms using confusion matrix,” Jurnal Sisfokom (Sistem Informasi dan Komputer), vol. 13, no. 1, pp. 92–97, 2024.

[12] S. B. Kotsiantis, I. D. Zaharakis, and P. E. Pintelas, “Machine learning: A review of classification and combining techniques,” Artificial Intelligence Review, vol. 26, no. 3, pp. 159–190, 2006.

[13] J. Alzubi, A. Nayyar, and A. Kumar, “Machine learning from theory to algorithms: An overview,” Journal of Physics: Conference Series, vol. 1142, no. 1, Art. no. 012012, 2018.

[14] R. Longadge and S. Dongre, “Class imbalance problem in data mining review,” arXiv preprint arXiv:1305.1707, 2013.

[15] D. Devi, S. K. Biswas, and B. Purkayastha, “A review on solution to class imbalance problem: Undersampling approaches,” in Proc. 2020 Int. Conf. Computational Performance Evaluation (ComPE), 2020, pp. 626–631.

[16] G. N. Ahmad, H. Fatima, S. Ullah, and A. S. Saidi, “Efficient medical diagnosis of human heart diseases using machine learning techniques with and without GridSearchCV,” IEEE Access, vol. 10, pp. 80151–80173, 2022.

[17] W. Aprilliandhika and F. F. Abdulloh, “Comparison of K-nearest neighbor and support vector machine algorithm optimization with grid search CV on stroke prediction,” Jurnal Teknik Informatika (Jutif), vol. 5, no. 4, pp. 991–1000, 2024.

[18] G. A. Pradipta, R. Wardoyo, A. Musdholifah, I. N. H. Sanjaya, and M. Ismail, “SMOTE for handling imbalanced data problem: A review,” in Proc. 2021 6th Int. Conf. Informatics and Computing (ICIC), 2021, pp. 1–8.

Downloads

Published

07-09-2026

How to Cite

Chauhan, D., Sharma, R., Shyam Ji, & Patwal, S. S. (2026). An AI-Driven Framework for Diabetes Prediction and Classification Using Medical Data, SMOTE, and Grid Search CV Optimization. Asian Journal of Engineering and Applied Technology, 15(2), 1–6. https://doi.org/10.70112/ajeat-2026.15.2.4370

Issue

Section

Research Article

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 > >> 

You may also start an advanced similarity search for this article.