Modifikasi Gain Ratio Pada Algoritma C4.5 dengan Nilai Koefisien Determinasi untuk Prediksi Kelulusan Mahasiswa (Studi Kasus: Universitas Islam Madura)
Abstract
C4.5 is a decision tree algorithm that can be used for making predictions. The stages start from forming a decision tree through splitting attributes, pruning and extracting rules or knowledge to then be used for prediction. However, one of the weaknesses of the C4.5 algorithm is the occurrence of overfitting and misclassification costs which result in low prediction performance. The development of the C4.5 algorithm has been carried out in terms of split attributes such as the imprecise info-gain ratio (Credal-C4.5) method using Imprecise Probability Theory, bossing gain ratio (C5.0) and average gain. This research applies the termination coefficient value (R2) as a method for modifying the gain ratio in selecting attributes as decision tree nodes which is then implemented to predict student graduation on time using a case study at the Universitas Islam Madura (UIM). Testing of the decision tree model rule for predicting student graduation on time at UIM shows that the performance values of accuracy, precision and recall are 70.49%, 77.14% and 72.97%. This performance is higher compared to the C4.5 algorithm without making modifications to the coefficient of determination, especially in accuracy and recall performance, while the precision is lower but the difference is below 1%. The difference in performance values was 11.48% (positive) for accuracy and 27.03% (positive) for recall. Meanwhile, precision performance has a difference of -0.13% (negative). The application of the Knowledge Model Rule for student graduation on time at SIMAT UIM shows very good results because it displays a prediction results page.
Keywords
References
E. S. Rahayu, R. Satria, and C. Supriyanto, “Penerapan Metode Average Gain, Threshold Pruning dan Cost Complexity Pruning Untuk Split Atribut Pada Algoritma C4.5,” Journal of Intelligent Systems, vol. 1, no. 2, pp. 91–97, 2015.
M. Bansal, A. Goyal, and A. Choudhary, “A comparative analysis of K-Nearest Neighbor, Genetic, Support Vector Machine, Decision Tree, and Long Short Term Memory algorithms in machine learning,” Decision Analytics Journal, vol. 3, p. 100071, Jun. 2022, doi: 10.1016/j.dajour.2022.100071.
M. Muhsi, “Model dan Analisa Faktor Eksternal Aktifitas Siswa Kelas X TKJ SMKN 1 Pakong Pamekasan Menggunakan Algoritma Decision Tree,” Jurnal Aplikasi Teknologi Informasi dan Manajemen (JATIM), vol. 2, no. 2, pp. 92–106, 2021, doi: 10.31102/jatim.v2i2.1239.
H. Bin Wang and Y. J. Gao, “Research on C4.5 algorithm improvement strategy based on MapReduce,” in Procedia Computer Science, Elsevier B.V., 2021, pp. 160–165. doi: 10.1016/j.procs.2021.02.045.
J. Wang, “Application of C4.5 Decision Tree Algorithm for Evaluating the College Music Education,” Mobile Information Systems, vol. 2022, 2022, doi: 10.1155/2022/7442352.
S. L. Salzberg, “C4.5: Programs for Machine Learning by J. Ross Quinlan. Morgan Kaufmann Publishers, Inc., 1993,” Mach Learn, vol. 16, no. 3, pp. 235–240, 1994, doi: 10.1007/bf00993309.
J. R. Quinlan, {C4}.5 - Programs for Machine Learning. San Mateo, California: Morgan Kaufmann Publishers, 1993.
C. J. Mantas and J. Abellán, “Credal-C4.5: Decision tree based on imprecise probabilities to classify noisy data,” Expert Syst Appl, vol. 41, no. 10, pp. 4625–4637, Aug. 2014, doi: 10.1016/j.eswa.2014.01.017.
G. Sarailidis, T. Wagener, and F. Pianosi, “Integrating scientific knowledge into machine learning using interactive decision trees,” Comput Geosci, vol. 170, Jan. 2023, doi: 10.1016/j.cageo.2022.105248.
N. V Chawla, “Many Are Better Than One: Improving Probabilistic Estimates from Decision Trees,” in Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment, J. Quiñonero-Candela, I. Dagan, B. Magnini, and F. d’Alché-Buc, Eds., Berlin, Heidelberg: Springer Berlin Heidelberg, 2006, pp. 41–55.
M. Muhsi, S. Suprapto, and R. Rofiuddin, “Node Selection Method for Split Attribute in C4.5 Algorithm Using the Coefficient of Determination Values for Multivariate Data Set,” Jurnal Penelitian Pendidikan IPA, vol. 9, no. 7, pp. 5574–5583, Jul. 2023, doi: 10.29303/jppipa.v9i7.4031.
A. Yalçınkaya, İ. G. Balay, and B. Şenoǧlu, “A new approach using the genetic algorithm for parameter estimation in multiple linear regression with long-tailed symmetric distributed error terms: An application to the Covid-19 data,” Chemometrics and Intelligent Laboratory Systems, vol. 216, Sep. 2021, doi: 10.1016/j.chemolab.2021.104372.
A. H. AL-Marshadi, M. Aslam, and A. Alharbey, “Selecting the covariance structure for the seemingly unrelated regression models,” J King Saud Univ Sci, vol. 34, no. 4, Jun. 2022, doi: 10.1016/j.jksus.2022.102027.
I. D. Mienye, Y. Sun, and Z. Wang, “Prediction performance of improved decision tree-based algorithms: A review,” in Procedia Manufacturing, Elsevier B.V., 2019, pp. 698–703. doi: 10.1016/j.promfg.2019.06.011.
J. R. Saura, D. Palacios-Marqués, and D. Ribeiro-Soriano, “Using data mining techniques to explore security issues in smart living environments in Twitter,” Comput Commun, vol. 179, pp. 285–295, Nov. 2021, doi: 10.1016/j.comcom.2021.08.021.
H. Thakkar, V. Shah, H. Yagnik, and M. Shah, “Comparative anatomization of data mining and fuzzy logic techniques used in diabetes prognosis,” Clinical eHealth, vol. 4, pp. 12–23, 2021, doi: 10.1016/j.ceh.2020.11.001.
J. Kalezhi, M. Chibuluma, C. Chembe, V. Chama, F. Lungo, and D. Kunda, “Modelling Covid-19 infections in Zambia using data mining techniques,” Results in Engineering, vol. 13, Mar. 2022, doi: 10.1016/j.rineng.2022.100363.
M. Patrício et al., “Using Resistin, glucose, age and BMI to predict the presence of breast cancer,” BMC Cancer, vol. 18, no. 1, 2018, doi: 10.1186/s12885-017-3877-1.
H. Sastypratiwi, Y. Yulianti, and H. Muhardi, “Uji Komparasi Algoritma Naïve Bayes dan Decision Tree Classification Menggunakan Covid-19 Dataset,” Jurnal Edukasi dan Penelitian Informatika (JEPIN), vol. 8, no. 1, p. 1, 2022, doi: 10.26418/jp.v8i1.49841.
T. Kristóf and M. Virág, “EU-27 bank failure prediction with C5.0 decision trees and deep learning neural networks,” Res Int Bus Finance, vol. 61, Oct. 2022, doi: 10.1016/j.ribaf.2022.101644.
M. S. Abubakari and S. Suprapto, “Educational Data Mining to Predict Students Performance Based on Deep Learning Neural Network,” in Proceeding International Conference On Health, Social Sciences And Technology, 2021, pp. 13–16. Accessed: Aug. 03, 2022. [Online]. Available: https://ojs.poltekkespalembang.ac.id/index.php/icohsst/article/view/697
I. G. T. Isa and F. Elfaladonna, “Penilaian Kinerja Akurasi Metode Klasifikasi dalam Dataset Penerimaan Mahasiswa Baru Universitas XYZ,” Jurnal Edukasi dan Penelitian Informatika (JEPIN), vol. 8, no. 2, p. 292, 2022, doi: 10.26418/jp.v8i2.54316.
A. Ermillian and K. Nugroho, “Perancangan Model Deteksi Potensi Siswa Putus Sekolah Menggunakan Metode Logistic Regression Dan Decision Tree,” Jurnal Informatika: Jurnal Pengembangan IT, vol. 9, no. 3, pp. 281–295, Dec. 2024, doi: 10.30591/jpit.v9i3.8007.
DOI: https://doi.org/10.30591/jpit.v10i1.5986
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution 4.0 International License.
JPIT INDEXED BY
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | |

This work is licensed under a Creative Commons Attribution 4.0 International License.









