Model Hybrid PSO-SMOTE, Correlation Feature Selection dan Random Forest untuk Deteksi Penyakit Jantung

Nur Dila Yuanti, Taghfirul Azhima Yoga Siswa, Wawan Joko Pranoto

Abstract


Heart disease is one of the leading causes of death worldwide, with approximately 17.8 million deaths reported in 2021. In Indonesia, the number of cases reached an estimated 15.5 million in 2022, highlighting the need for accurate early detection. This study aims to improve heart disease classification by integrating Correlation Feature Selection (CFS), Synthetic Minority Oversampling Technique (SMOTE), Particle Swarm Optimization (PSO), and Random Forest. CFS was applied to remove less relevant features before classification, SMOTE addressed class imbalance, and PSO optimized the Random Forest hyperparameters. The dataset consisted of 1,025 records from four heart disease datasets, which were cleaned into 302 unique instances. CFS eliminated the fasting blood sugar (fbs) attribute due to its very weak correlation with the target variable. Evaluation using 10-fold cross-validation showed that the baseline Random Forest achieved an accuracy of 82.80%, precision of 83.62%, recall of 86.07%, and f1-score of 84.46%, while Random Forest–PSO with SMOTE produced the best performance, achieving 85.11% accuracy, 84.62% precision, 89.74% recall, and 86.77% f1-score. These findings indicate that PSO provides the greatest performance improvement, while CFS simplifies the feature set before classification. The proposed hybrid model contributes by integrating feature selection, data balancing, and hyperparameter optimization into a single classification framework for more effective heart disease prediction.

Keywords


Correlation, Heart Disease, Particle Swarm Optimization, Random Forest, SMOTE

Full Text:

References


K. K. R. Indonesia, “Kementerian Kesehatan 2024,” 2024, LMS Kementerian Kesehatan RI. [Online]. Available: https://lms.kemkes.go.id/courses/35bff824-437e-4557-b37a-94b128c43333

A. H. Anwar, “Sistematic Review Faktor Resiko Penyakit Jantung,” Indones. J. Heal. Res. Innov., vol. 02, no. 01, pp. 57–69, 2025, doi: https://doi.org/10.64094/fqanc998.

G. A. Ansari, S. S. Bhat, M. D. Ansari, S. Ahmad, J. Nazeer, and A. E. M. Eljialy, “Performance Evaluation of Machine Learning Techniques (MLT) for Heart Disease Prediction,” Comput. Math. Methods Med., vol. 2023, no. Ml, 2023, doi: 10.1155/2023/8191261.

V. Lampos, J. Mintz, and X. Qu, “An artificial intelligence approach for selecting effective teacher communication strategies in autism education,” npj Sci. Learn., vol. 6, no. 1, 2021, doi: 10.1038/s41539-021-00102-x.

E. Erlin, Y. Desnelita, N. Nasution, L. Suryati, and F. Zoromi, “Dampak SMOTE terhadap Kinerja Random Forest Classifier berdasarkan Data Tidak seimbang,” MATRIK J. Manajemen, Tek. Inform. dan Rekayasa Komput., vol. 21, no. 3, pp. 677–690, 2022, doi: 10.30812/matrik.v21i3.1726.

P. A. Jusia, A. Rahim, H. Yani, and J. Jasmir, “Improving Performance of KNN and C4.5 using Particle Swarm Optimization in Classification of Heart Diseases,” J. RESTI, vol. 8, no. 3, pp. 333–339, 2024, doi: 10.29207/resti.v8i3.5710.

N. Sureja, B. Chawda, and A. Vasant, “A novel salp swarm clustering algorithm for prediction of the heart diseases,” Indones. J. Electr. Eng. Comput. Sci., vol. 25, no. 1, pp. 265–272, 2022, doi: 10.11591/ijeecs.v25.i1.pp265-272.

A. Khaleel Faieq and M. M. Mijwil, “Prediction of heart diseases utilising support vector machine and artificial neural network,” Indones. J. Electr. Eng. Comput. Sci., vol. 26, no. 1, pp. 374–380, 2022, doi: 10.11591/ijeecs.v26.i1.pp374-380.

M. H. Musyaffa, T. H. Saragih, D. T. Nugrahadi, D. Kartini, and A. Farmadi, “Effectiveness of SMOTE in Enhancing Adult Autism Spectrum Disorder Diagnosis Predictive Performance With Missforest Imputation And Random Forest,” Indones. J. Electron. Electromed. Eng. Med. Informatics, vol. 7, no. 2, pp. 270–280, 2025, doi: 10.35882/ijeeemi.v7i2.66.

H. A. Salman, A. Kalakech, and A. Steiti, “Random Forest Algorithm Overview,” Babylonian J. Mach. Learn., vol. 2024, pp. 69–79, 2024, doi: 10.58496/bjml/2024/007.

A. Samosir, M. S. Hasibuan, W. E. Justino, and T. Hariyono, “Komparasi Algoritma Random Forest, Naïve Bayes dan K- Nearest Neighbor Dalam klasifikasi Data Penyakit Jantung,” Pros. Semin. Nas. Darmajaya, vol. 1, no. 0, pp. 214–222, 2021, [Online]. Available: https://jurnal.darmajaya.ac.id/index.php/PSND/article/view/2955

A. P. Siregar, D. P. Purba, J. P. Pasaribu, and K. R. Bakara, “Implementasi Algoritma Random Forest Dalam Klasifikasi Diagnosis Penyakit Stroke,” J. Penelit. Rumpun Ilmu Tek., vol. 2, no. 4, pp. 155–164, 2023.

A. Ghozali, H. Pratiwi, and S. S. Handajani, “Implementasi Data Mining Menggunakan Metode Random Forest Dan Support Vector Machine Dalam Klasifikasi Penyakit Diabetes,” Delta J. Ilm. Pendidik. Mat., vol. 11, no. 2, p. 147, 2023, doi: 10.31941/delta.v11i2.2686.

Agung Khoeruddin, Fahri Andriansyah Sudrajat, Galuh Purnama, Iman Kuwangid, Kurnia Kurnia, and Ricky Firmansyah, “Optimasi Fitur Seleksi Random Forest Menggunakan GA Dalam Klasifikasi Data Penyakit Gagal Jantung,” J. Penelit. Teknol. Inf. dan Sains, vol. 1, no. 2, pp. 01–09, 2023, doi: 10.54066/jptis.v1i2.323.

M. Hasan, M. A. Sahid, M. P. Uddin, M. A. Marjan, S. Kadry, and J. Kim, “Performance discrepancy mitigation in heart disease prediction for multisensory inter-datasets,” PeerJ Comput. Sci., vol. 10, pp. 1–51, 2024, doi: 10.7717/peerj-cs.1917.

Z. Noroozi, A. Orooji, and L. Erfannia, “Analyzing the impact of feature selection methods on machine learning algorithms for heart disease prediction,” Sci. Rep., no. 0123456789, pp. 1–15, 2023, doi: 10.1038/s41598-023-49962-w.

M. Salman, A. Nag, M. Mohisn, and S. Dev, “Healthcare Analytics Analyzing the impact of feature selection on the accuracy of heart disease prediction,” Healthc. Anal., vol. 2, no. February, p. 100060, 2022, doi: 10.1016/j.health.2022.100060.

H. Optimization, K. Vishnu, V. Reddy, I. Elamvazuthi, A. A. Aziz, and S. Paramasivam, “applied sciences An Efficient Prediction System for Coronary Heart Disease Risk Using Selected Principal Components and,” 2023.

U. Hasanah, A. M. Soleh, and K. Sadik, “Effect of Random Under sampling, Oversampling, and SMOTE on the Performance of Cardiovascular Disease Prediction Models,” J. Mat. Stat. dan Komputasi, vol. 21, no. 1, pp. 88–102, 2024, doi: 10.20956/j.v21i1.35552.

U. Ungkawa and M. A. Rafi, “Data Balancing Techniques Using the PCA-KMeans and ADASYN for Possible Stroke Disease Cases,” J. Online Inform., vol. 9, no. 1, pp. 138–147, 2024, doi: 10.15575/join.v9i1.1293.

A. A. Dhani, T. A. Y. Siswa, and W. J. Pranoto, “Perbaikan Akurasi Random Forest Dengan ANOVA Dan SMOTE Pada Klasifikasi Data Stunting,” Teknika, vol. 13, no. 2, pp. 264–272, 2024, doi: 10.34148/teknika.v13i2.875.

A. Singh, N. Prakash, and A. Jain, “Particle Swarm Optimization-Based Random Forest Framework for the Classification of Chronic Diseases,” IEEE Access, vol. 11, no. December, pp. 133931–133946, 2023, doi: 10.1109/ACCESS.2023.3335314.

R. Z. Zheng, Y. Zhang, and K. Yang, “A transfer learning-based particle swarm optimization algorithm for travelling salesman problem,” J. Comput. Des. Eng., vol. 9, no. 3, pp. 933–948, 2022, doi: 10.1093/jcde/qwac039.

K. A. Barry, Y. Manzali, R. Flouchi, and M. Elfar, “Heart disease approach using modified random forest and particle swarm optimization,” IAES Int. J. Artif. Intell., vol. 14, no. 2, pp. 1242–1251, 2025, doi: 10.11591/ijai.v14.i2.pp1242-1251.

T. P. Amal Dluha, Rizkianda Farma Saputra, Santoso, “Sentiment Analysis Aplication Ruang Guru Using Naive Bayes and Support Vector Machine,” JAISEN JournalJournal Adv. Inf. Syst. Eng., vol. 1, pp. 39–47, 2025, [Online]. Available: https://journal.unilak.ac.id/index.php/jaisen/home

F. Aldi et al., “STANDARDSCALER ’ S POTENTIAL IN ENHANCING BREAST CANCER,” vol. 5, no. 1, pp. 401–413, 2023.

C. Fan, M. Chen, X. Wang, J. Wang, and B. Huang, “A Review on Data Preprocessing Techniques Toward Ef fi cient and Reliable Knowledge Discovery From Building Operational Data,” vol. 9, no. March, pp. 1–17, 2021, doi: 10.3389/fenrg.2021.652801.

S. Akinwamide, T. Fele, and O. A. Ojo, “Comparative evaluation of supervised machine learning algorithms for breast cancer prediction using the Wisconsin diagnostic dataset,” vol. 24, no. 02, pp. 196–203, 2025.

N. S. Sani, M. I. Esa, and B. A. Musawi, “SS symmetry Feature Selection,” Symmetry (Basel)., vol. 15, no. 1, pp. 1–21, 2023, doi: 10.3390/sym15010123.

V. Artanti, “Classification of Cardiovascular Diseases Using the K-Nearest Neighbors (KNN) Algorithm,” IEESE Int. J. Sci. Technol., vol. 13, no. 2, pp. 12–27, 2024, doi: https://doi.org/10.62411/tc.v23i2.10061.

R. R. Adhitya, Wina Witanti, and Rezki Yuniarti, “Perbandingan Metode Cart Dan Naïve Bayes Untuk Klasifikasi Customer Churn,” INFOTECH J., vol. 9, no. 2, pp. 307–318, 2023, doi: 10.31949/infotech.v9i2.5641.

I. K. Nti, “Performance of Machine Learning Algorithms with Different K Values in K-fold Cross- Validation,” I.J. Inf. Technol. Comput. Sci., no. December, pp. 61–71, 2021, doi: 10.5815/ijitcs.2021.06.05.

D. H. Depari, Y. Widiastiwi, and M. M. Santoni, “Perbandingan Model Decision Tree, Naive Bayes dan Random Forest untuk Prediksi Klasifikasi Penyakit Jantung,” Inform. J. Ilmu Komput., vol. 18, no. 3, p. 239, 2022, doi: https://doi.org/10.52958/iftk.v18i3.4694.

S. Sathyanarayanan, “Confusion Matrix-Based Performance Evaluation Metrics,” African J. Biomed. Res., no. November, pp. 4023–4031, 2024, doi: 10.53555/ajbr.v27i4s.4345.

D. Kurniadi, F. Nuraeni, N. Faturrohman, and A. Mulyani, “Klasifikasi Perputaran Karyawan Perusahaan Menggunakan Algoritma Random Forest dan Random Over-sampling,” Edu Komputika J., vol. 10, no. 2, pp. 93–103, 2024, doi: 10.15294/edukomputika.v10i2.73782.

A. E. Saputra Rusdi, A. Rangkuti, N. Aris et., “Perbandingan ANN, Random Forest, dan XGBOOST Dalam Klasifikasi Antibiotik Dengan Penerapan Metode Sampling,” vol. 12, no. 4, 2025.

Z. Fan, B. Liu, and X. Yan, “Cardiovascular Disease Prediction Based on Machine Learning,” pp. 404–411, 2024, doi: 10.5220/0012939000004508.

W. J. P. Azwar Damari, Taghfirul Azhima Yoga Siswa, “IMPLEMENTATION OF THE PSO-SMOTE METHOD ON THE NAIVE BAYES ALGORITHM TO ADDRESS CLASS IMBALANCE IN LANDSLIDE DISASTER DATA PENERAPAN METODE PSO-SMOTE PADA ALGORITMA NAIVE BAYES UNTUK MENGATASI CLASS IMBALANCE,” vol. 10, no. 1, pp. 332–343, 2025.

M. Carvalho, A. J. Pinho, and S. Brás, “Resampling approaches to handle class imbalance : a review from a data perspective,” J. Big Data, vol. 55, pp. 1–29, 2025, doi: 10.1186/s40537-025-01119-4.

A. Arafa, N. El-fishawy, M. Badawy, and M. Radad, “RN-SMOTE : Reduced Noise SMOTE based on DBSCAN for enhancing imbalanced data classification,” J. King Saud Univ. - Comput. Inf. Sci., vol. 34, no. 8, pp. 5059–5074, 2022, doi: 10.1016/j.jksuci.2022.06.005.

A. Sakho and E. Malherbe, “Do we need rebalancing strategies ? A theoretical and empirical study around SMOTE and its variants,” arXiv Prepr., pp. 1–18, 2025, doi: https://arxiv.org/abs/2402.03819.




DOI: https://doi.org/10.30591/jpit.v11i2.10241

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.