Optimization of Random Forest Algorithm Performance for Early Detection of Stroke Disease Using Medical Record Data

Theo Krisna Amarya, Aidina Ristyawan, Rina Firliana

Abstract


Stroke is a medical condition that occurs when blood flow to the brain is blocked, causing damage to brain tissue. Stroke is the second largest cause of death and disability in the world, this disease can affect all ages and is influenced by various risk aspects, such as unhealthy lifestyles, high blood pressure, high blood sugar levels, and other risks. It is very important to detect stroke in patients as soon as possible to prevent it. This study proposes the optimization of the performance of the Random Forest algorithm as an early detection model for stroke by utilizing a hybrid sampling method called SMOTETomek and also conducting several experiments on the parameter settings of the Random Forest algorithm. The results of this study show an increase compared to the previous one which had an accuracy was 94% with a standard deviation of 2%, In this study, it managed to reach accuracy of 96% with a standard deviation of 0% with a ROC curve (AUC) value of 0.96 or 96%. The algorithm that has 96% accuracy in the discussion is Random Forest Algorithm as estimator of AdaBoost.

Keywords


Adaboost;Hybrid Sampling;Random Forest;Smote-Tomek;Stroke Predict

Full Text:

References


W. Riyadina and E. Rahajeng, “Determinan Penyakit Stroke,” Kesmas Natl. Public Heal. J., vol. 7, no. 7, p. 324, Feb. 2013, doi: 10.21109/kesmas.v7i7.31.

A. Ramadhanu, R. Ayu Mahessya, M. Raihan Zaky, M. Isra, S. Informasi, and U. Putra Indonesia YPTK Padang, “PENERAPAN TEKNOLOGI MACHINE LEARNING DENGAN METODE VADER PADA APLIKASI SENTIMEN TAMU DI HOTEL DYMENS,” JOISIE J. Inf. Syst. Informatics Eng., vol. 7, no. 1, pp. 165–173, Jun. 2023.

M. S. Pathan, Z. Jianbiao, D. John, A. Nag, and S. Dev, “Identifying Stroke Indicators Using Rough Sets,” IEEE Access, vol. 8, pp. 210318–210327, Nov. 2020, doi: 10.1109/ACCESS.2020.3039439.

A. P. Wibawa, M. Guntur, A. Purnama, M. Fathony Akbar, and F. A. Dwiyanto, “Metode-metode Klasifikasi,” Pros. Semin. Ilmu Komput. dan Teknol. Inf., vol. 3, no. 1, 2018.

U. Bradter, J. D. Altringham, W. E. Kunin, T. J. Thom, J. O’Connell, and T. G. Benton, “Variable ranking and selection with random forest for unbalanced data,” Environ. Data Sci., vol. 1, pp. 1–23, Nov. 2022, doi: 10.1017/eds.2022.34.

T. Krisna Amarya, A. G. Candra Andy, R. Achmad, E. Daniati, and A. Ristyawan, “Analisa Perbandingan Algoritma Classification Berdasarkan Komposisi Label,” Pros. SEMNAS INOTEK, vol. 8, pp. 2549–7952, Aug. 2024, Accessed: Sep. 30, 2024. [Online]. Available: https://proceeding.unpkediri.ac.id/index.php/inotek

M. Guhdar, A. Ismail Melhum, and A. Luqman Ibrahim, “Optimizing Accuracy of Stroke Prediction Using Logistic Regression,” J. Technol. Informatics, vol. 4, no. 2, pp. 41–47, 2023, doi: 10.37802/joti.v4i2.278.

E. Wulandari et al., “Classification Of Stroke Prediction Using The Support Vector Machine (SVM) Method 1,2,” J. Tek. Inform. dan Sist. Inf., vol. 11, no. 3, 2024.

W. Wang et al., “A systematic review of machine learning models for predicting outcomes of stroke with structured data,” PLoS One, vol. 15, no. 6, pp. 1–16, 2020, doi: 10.1371/journal.pone.0234722.

X. Huang et al., “Novel Insights on Establishing Machine Learning-Based Stroke Prediction Models Among Hypertensive Adults,” Front. Cardiovasc. Med., vol. 9, no. May, pp. 1–11, 2022, doi: 10.3389/fcvm.2022.901240.

Muvida and F. Amar, “The Features of Comorbidity of Stroke in The Indonesian Population: Findings from The Indonesian Family Life Survey (IFLS-5),” Magna Neurol., vol. 2, no. 2, pp. 42–47, Jul. 2024, doi: 10.20961/magnaneurologica.v2i2.948.

M. Guhdar, A. Ismail Melhum, and A. Luqman Ibrahim, “Optimizing Accuracy of Stroke Prediction Using Logistic Regression,” J. Technol. Informatics, vol. 4, no. 2, pp. 41–47, Apr. 2023, doi: 10.37802/joti.v4i2.278.

E. Wulandari et al., “Classification Of Stroke Prediction Using The Support Vector Machine (SVM),” J. Tek. Inform. dan Sist. Inf., vol. 11, no. 3, pp. 17–29, Sep. 2024.

T. K. Amarya, A. C. A. G, R. Achmad, E. Daniati, and A. Ristyawan, “Analisa Perbandingan Algoritma Classification Berdasarkan Komposisi Label,” Semin. Nas. Inov. Teknol., vol. 8, pp. 32–40, 2024.

Fedesoriano, “Stroke Prediction Dataset,” https://www.kaggle.com/, 2024. https://www.kaggle.com/datasets/fedesoriano/stroke-prediction-dataset/data (accessed Oct. 28, 2024).

Y. Ye, Q. Wu, J. Zhexue Huang, M. K. Ng, and X. Li, “Stratified sampling for feature subspace selection in random forests for high dimensional data,” Pattern Recognit., vol. 46, no. 3, pp. 769–787, Mar. 2013, doi: 10.1016/J.PATCOG.2012.09.005.

H. Fei et al., “Cotton Classification Method at the County Scale Based on Multi-Features and Random Forest Feature Selection Algorithm and Classifier,” Remote Sens., vol. 14, no. 4, pp. 1–28, Feb. 2022, doi: 10.3390/rs14040829.

B. H. Sadiq and S. R. Zeebaree, “Parallel Processing Impact on Random Forest Classifier Performance: A CIFAR-10 Dataset Study,” Indones. J. Comput. Sci. Attrib., vol. 13, no. 2, pp. 1833–1846, Apr. 2024.

M. Shahhosseini and G. Hu, “Improved Weighted Random Forest for Classification Problems.”

A. J. Wyner, M. Olson, J. Bleich, and D. Mease, “Explaining the Success of AdaBoost and Random Forests as Interpolating Classifiers,” 2017. [Online]. Available: http://jmlr.org/papers/v18/15-240.html.




DOI: https://doi.org/10.30591/jpit.v10i3.8424

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.