Brain Stroke Prediction Using Artificial Intelligence and Machine Learning; A Comprehensive Risk Assessment Model
Abstract
Brain stroke refers to a serious condition, happens when the blood flow to the brain is either blocked or affected owing to bleeding due to raptured vessel, in the brain. This lack of blood, thus lack of oxygen supply to brain cells harms brain tissue and leads to quick cell death. Each year, about fifteen million people around the world experienced the stroke. Of these, five million die, and another five million survive but face lasting disabilities. The mortality rate within 28 days after a stroke is about 28%. This study aims to develop a predictive model using artificial intelligence (AI) algorithms for assessing individual stroke risk. It uses a publicly available dataset of about 5,000 records from a website called Kaggle. The predictors include demographic and clinical variables including gender, hypertension, heart disease, smoking status etc. Seven machine learning classifiers, namely Decision Tree, Random Forest Classifier, Support Vector Machine SVM, Naïve Bayes, Logistic Regression, K-Nearest Neighbor KNN, and Gradient Boosting is implemented in Python. These classifiers are compared using different evaluation parameters. The result reveals that the Random Forest classifier consistently outperformed all other models in the evaluation parameters, achieving the highest predictive performance for stroke risk assessment.
References
[2] Pandian, J. D., Gall, S. L., Kate, M. P., Silva, G. S., Akinyemi, R. O., Ovbiagele, B. I., ... & Thrift, A. G. (2018). Prevention of stroke: a global perspective. The Lancet, 392(10154), 1269-1278.
[3] George, M. G., Tong, X., & Bowman, B. A. (2017). Prevalence of cardiovascular risk factors and strokes in younger adults. JAMA neurology, 74(6), 695-703.
[4] Sacco, R. L., Benjamin, E. J., Broderick, J. P., Dyken, M., Easton, J. D., Feinberg, W. M., ... & Wolf, P. A. (1997). Risk factors. Stroke, 28(7), 1507-1517.
[5] Strong, K., Mathers, C., & Bonita, R. (2007). Preventing stroke: saving lives around the world. The Lancet Neurology, 6(2), 182-187.
[6] Lawes, C. M., Bennett, D. A., Feigin, V. L., & Rodgers, A. (2004). Blood pressure and stroke: an overview of published reviews. Stroke, 35(3), 776-785.
[7] Mukherjee, D., & Patil, C. G. (2011). Epidemiology and the global burden of stroke. World neurosurgery, 76(6), S85-S90.
[8] Schroeder, E. B., Rosamond, W. D., Morris, D. L., Evenson, K. R., & Hinn, A. R. (2000). Determinants of use of emergency medical services in a population with stroke symptoms: the Second Delay in Accessing Stroke Healthcare (DASH II) Study. Stroke, 31(11), 2591-2596.
[9] Sirsat, M. S., Fermé, E., & Câmara, J. (2020). Machine learning for brain stroke: a review. Journal of Stroke and Cerebrovascular Diseases, 29(10), 105162.
[10] Amarenco, P., Bogousslavsky, J., Caplan, L. R., Donnan, G. A., & Hennerici, M. G. (2009). Classification of stroke subtypes. Cerebrovascular diseases, 27(5), 493-501.
[11] Boehme, A. K., Esenwa, C., & Elkind, M. S. (2017). Stroke risk factors, genetics, and prevention. Circulation research, 120(3), 472-495.
[12] Aigner, A., Grittner, U., Rolfs, A., Norrving, B., Siegerink, B., & Busch, M. A. (2017). Contribution of established stroke risk factors to the burden of stroke in young adults. Stroke, 48(7), 1744-1751.
[13] Robertson, I. H., & Murre, J. M. (1999). Rehabilitation of brain damage: brain plasticity and principles of guided recovery. Psychological bulletin, 125(5), 544.
[14] Bonkhoff, A. K., & Grefkes, C. (2022). Precision medicine in stroke: towards personalized outcome predictions using artificial intelligence. Brain, 145(2), 457-475.
[15] Mainali, S., Darsie, M. E., & Smetana, K. S. (2021). Machine learning in action: stroke diagnosis and outcome prediction. Frontiers in neurology, 12, 734345.
[16] Feske, S. K. (2021). Ischemic stroke. The American journal of medicine, 134(12), 1457-1464.
[17] Herpich, F., & Rincon, F. (2020). Management of acute ischemic stroke. Critical care medicine, 48(11), 1654-1663.
[18] Caplan, L. R. (1993). Brain embolism, revisited. Neurology, 43(7), 1281-1281.
[19] National Institute of Neurological Disorders and Stroke rt-PA Stroke Study Group. (1995). Tissue plasminogen activator for acute ischemic stroke. New England Journal of Medicine, 333(24), 1581-1588.
[20] Montaño, A., Hanley, D. F., & Hemphill III, J. C. (2021). Hemorrhagic stroke. Handbook of clinical neurology, 176, 229-248.
[21] Teunissen, L. L., Rinkel, G. J., Algra, A., & Van Gijn, J. (1996). Risk factors for subarachnoid hemorrhage: a systematic review. Stroke, 27(3), 544-549.
[22] Wolf, P. A., D'Agostino, R. B., Belanger, A. J., & Kannel, W. B. (1991). Probability of stroke: a risk profile from the Framingham Study. Stroke, 22(3), 312-318.
[23] Makke, N., & Chawla, S. (2024). Interpretable scientific discovery with symbolic regression: a review. Artificial Intelligence Review, 57(1), 2.
[24] Rajula, H. S. R., Verlato, G., Manchia, M., Antonucci, N., & Fanos, V. (2020). Comparison of conventional statistical methods with machine learning in medicine: diagnosis, drug development, and treatment. Medicina, 56(9), 455.
[25] Rimal, Y., Sharma, N., Paudel, S., Alsadoon, A., Koirala, M. P., & Gill, S. (2025). Comparative analysis of heart disease prediction using logistic regression, SVM, KNN, and random forest with cross-validation for improved accuracy. Scientific Reports, 15(1), 13444.
[26] Emon, M. U., Keya, M. S., Meghla, T. I., Rahman, M. M., Al Mamun, M. S., & Kaiser, M. S. (2020, November). Performance analysis of machine learning approaches in stroke prediction. In 2020 4th international conference on electronics, communication and aerospace technology (ICECA) (pp. 1464-1469). IEEE.
[27] Dritsas, E., & Trigka, M. (2022). Stroke risk prediction with machine learning techniques. Sensors, 22(13), 4670.
[28] Biswas, N., Uddin, K. M. M., Rikta, S. T., & Dey, S. K. (2022). A comparative analysis of machine learning classifiers for stroke prediction: A predictive analytics approach. Healthcare Analytics, 2, 100116.
[29] Nielsen, A., Hansen, M. B., Tietze, A., & Mouridsen, K. (2018). Prediction of tissue outcome and assessment of treatment effect in acute ischemic stroke using deep learning. Stroke, 49(6), 1394-1401.
[30] Hung, C. Y., Chen, W. C., Lai, P. T., Lin, C. H., & Lee, C. C. (2017, July). Comparing deep neural network and other machine learning algorithms for stroke prediction in a large-scale population-based electronic medical claims database. In 2017 39th annual international conference of the IEEE engineering in medicine and biology society (EMBC) (pp. 3110-3113). IEEE.
[31] Mondal, S., Ghosh, S., & Nag, A. (2024). Brain stroke prediction model based on boosting and stacking ensemble approach. International Journal of Information Technology, 16(1), 437-446.
[32] Collins, G. S., Dhiman, P., Navarro, C. L. A., Ma, J., Hooft, L., Reitsma, J. B., ... & Moons, K. G. (2021). Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ open, 11(7), e048008.
[33] Bella, A., Ferri, C., Hernández-Orallo, J., & Ramírez-Quintana, M. J. (2010). Calibration of machine learning models. In Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques (pp. 128-146). IGI Global Scientific Publishing.
[34] Zimmerman, N., Presto, A. A., Kumar, S. P., Gu, J., Hauryliuk, A., Robinson, E. S., & Robinson, A. L. (2018). A machine learning calibration model using random forests to improve sensor performance for lower-cost air quality monitoring. Atmospheric Measurement Techniques, 11(1), 291-313.
[35] Huang, J., Galal, G., Etemadi, M., & Vaidyanathan, M. (2022). Evaluation and mitigation of racial bias in clinical machine learning models: scoping review. JMIR medical informatics, 10(5), e36388.
[36] Straw, I., & Wu, H. (2022). Investigating for bias in healthcare algorithms: a sex-stratified analysis of supervised machine learning models in liver disease prediction. BMJ health & care informatics, 29(1), e100457.
[37] Holmes, J., Sacchi, L., & Bellazzi, R. (2004). Artificial intelligence in medicine. Ann R Coll Surg Engl, 86(86), 334-8.
[38] Holzinger, A., Langs, G., Denk, H., Zatloukal, K., & Müller, H. (2019). Causability and explainability of artificial intelligence in medicine. Wiley interdisciplinary reviews: data mining and knowledge discovery, 9(4), e1312.
[39] Mukhamediev, R. I., Popova, Y., Kuchin, Y., Zaitseva, E., Kalimoldayev, A., Symagulov, A., ... & Yelis, M. (2022). Review of artificial intelligence and machine learning technologies: classification, restrictions, opportunities and challenges. Mathematics, 10(15), 2552.
[40] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. nature, 521(7553), 436-444.
[41] Pandey, D., Niwaria, K., & Chourasia, B. (2019). Machine learning algorithms: a review. Mach. Learn, 6(2).
[42] Sarker, I. H. (2021). Machine learning: Algorithms, real-world applications and research directions. SN computer science, 2(3), 160.
[43] Qamar, R., & Zardari, B. A. (2023). Artificial neural networks: An overview. Mesopotamian Journal of Computer Science, 2023, 124-133.
[44] Abdou, M. A. (2022). Literature review: Efficient deep neural networks techniques for medical image analysis. Neural Computing and Applications, 34(8), 5791-5812.
[45] Banerjee, I., Ling, Y., Chen, M. C., Hasan, S. A., Langlotz, C. P., Moradzadeh, N., ... & Lungren, M. P. (2019). Comparative effectiveness of convolutional neural network (CNN) and recurrent neural network (RNN) architectures for radiology text report classification. Artificial intelligence in medicine, 97, 79-88.
[46] Ng, A., & Jordan, M. (2001). On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. Advances in neural information processing systems, 14.
[47] Schein, A. I., & Ungar, L. H. (2007). Active learning for logistic regression: an evaluation. Machine Learning, 68(3), 235-265.
[48] LaValley, M. P. (2008). Logistic regression. Circulation, 117(18), 2395-2399.
[49] Safavian, S. R., & Landgrebe, D. (1991). A survey of decision tree classifier methodology. IEEE transactions on systems, man, and cybernetics, 21(3), 660-674.
[50] Priyanka, & Kumar, D. (2020). Decision tree classifier: a detailed survey. International Journal of Information and Decision Sciences, 12(3), 246-269.
[51] Pal, M. (2005). Random forest classifier for remote sensing classification. International journal of remote sensing, 26(1), 217-222.
[52] Belgiu, M., & Drăguţ, L. (2016). Random forest in remote sensing: A review of applications and future directions. ISPRS journal of photogrammetry and remote sensing, 114, 24-31.
[53] Bentéjac, C., Csörgő, A., & Martínez-Muñoz, G. (2021). A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review, 54(3), 1937-1967.
[54] Natekin, A., & Knoll, A. (2013). Gradient boosting machines, a tutorial. Frontiers in neurorobotics, 7, 21.
[55] Suthaharan, S. (2016). Support vector machine. In Machine learning models and algorithms for big data classification: thinking with examples for effective learning (pp. 207-235). Boston, MA: Springer US.
[56] Fung, G., & Mangasarian, O. L. (2001, August). Proximal support vector machine classifiers. In Proceedings of the seventh ACM SIGKDD international conference on Knowledge discovery and data mining (pp. 77-86).
[57] Peterson, L. E. (2009). K-nearest neighbor. Scholarpedia, 4(2), 1883.
[58] Liao, Y., & Vemuri, V. R. (2002). Use of k-nearest neighbor classifier for intrusion detection. Computers & security, 21(5), 439-448.
[59] Patil, T. R., & Sherekar, S. S. (2013). Performance analysis of Naive Bayes and J48 classification algorithm for data classification. International journal of computer science and applications, 6(2), 256-261.
[60] Dai, W., Xue, G. R., Yang, Q., & Yu, Y. (2007, July). Transferring naive bayes classifiers for text classification. In AAAI (Vol. 7, pp. 540-545).
[61] Kamiran, F., & Calders, T. (2012). Data preprocessing techniques for classification without discrimination. Knowledge and information systems, 33(1), 1-33.
[62] Chu, X., Ilyas, I. F., Krishnan, S., & Wang, J. (2016, June). Data cleaning: Overview and emerging challenges. In Proceedings of the 2016 international conference on management of data (pp. 2201-2206).
[63] Saar-Tsechansky, M., & Provost, F. (2007). Handling missing values when applying classification models.
[64] Foken, T., Leuning, R., Oncley, S. R., Mauder, M., & Aubinet, M. (2011). Corrections and data quality control. In Eddy covariance: a practical guide to measurement and data analysis (pp. 85-131). Dordrecht: Springer Netherlands.
[65] Singh, D., & Singh, B. (2020). Investigating the impact of data normalization on classification performance. Applied Soft Computing, 97, 105524.
[66] Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research, 16, 321-357.
[67] Raschka, S., Patterson, J., & Nolet, C. (2020). Machine learning in python: Main developments and technology trends in data science, machine learning, and artificial intelligence. Information, 11(4), 193.
[68] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., ... & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. the Journal of machine Learning research, 12, 2825-2830.
[69] Müller, A. C., & Guido, S. (2016). Introduction to machine learning with Python: a guide for data scientists. " O'Reilly Media, Inc.".
[70] Van Rossum, G., & Drake Jr, F. L. (1995). Python tutorial (Vol. 620). Amsterdam, The Netherlands: Centrum voor Wiskunde en Informatica.

