Optimasi Metode Ensemble Stacking Dalam Klasifikasi Kualitas Air Berbasis Algoritma XGBoost, LightGBM, dan CatBoost

Raihan Aly Zaky, Windha Mega Pradnya Dhuhita

Abstract


Kualitas air merupakan isu lingkungan yang berdampak langsung terhadap kesehatan manusia dan keberlanjutan ekosistem, sehingga diperlukan sistem deteksi kualitas air yang akurat, cepat, dan andal untuk mendukung pengambilan keputusan. Penelitian ini bertujuan mengoptimalkan kinerja klasifikasi kualitas air dengan menerapkan metode Ensemble Stacking berbasis algoritma XGBoost, LightGBM, dan CatBoost, dengan Random Forest sebagai meta-learner. Proses penelitian meliputi pemuatan dataset kualitas air, eksplorasi data, tahap preprocessing, pelatihan model dasar, serta pembentukan skema stacking. Untuk meningkatkan performa model, diterapkan mekanisme passthrough atau penambahan fitur asli yang dikombinasikan dengan optimasi hyperparameter menggunakan RandomizedSearchCV guna memperoleh konfigurasi model terbaik. Hasil pengujian menunjukkan bahwa metode Ensemble Stacking mampu mencapai akurasi sebesar 97,06%. Setelah dilakukan optimasi melalui passthrough dan tuning hyperparameter, performa model meningkat dengan akurasi tertinggi sebesar 97,12% dan nilai precision mencapai 90,48%. Model yang diusulkan terbukti mampu menurunkan kesalahan prediksi, khususnya False Positive, sehingga lebih aman dalam penentuan status kualitas air. Dengan demikian, penelitian ini menghasilkan model prediksi yang andal dan berpotensi diterapkan pada instansi lingkungan, fasilitas pengolahan air, serta sistem pemantauan kualitas air berbasis data.


Keywords


Kualitas Air; Klasifikasi; Ensemble Stacking; XGBoost; LightGBM; CatBoost; Random Forest

Full Text:

References


N. Akmal, M. Faez, A. Jalal, K. Chowdhury, and N. Hafizah, “Water pollution and the assessment of water

quality parameters : a review,” Desalin. Water Treat., vol. 294, pp. 79–88, 2023, doi: 10.5004/dwt.2023.29433.

I. Juwana, S. V. Lazuardi, and H. Herdiansyah, “Investigating the Factor of Water Consumption Regarding

the Impact and Implementation of Water Governance in Urban Areas,” vol. 7, no. 1, pp. 55–64, 2024.

W. P. Anggraini, D. I. Kusumastuti, E. P. Wahono, T. Sipil, and U. Lampung, “Sistem Informasi Kualitas Air

Sungai di Wilayah Sungai Seputih,” vol. 9, no. 4, pp. 964–978, 2024.

H. Zhang et al., “Natural and anthropogenic imprints on seasonal river water quality trends across China,”

, doi: 10.1038/s41545-025-00481-3.

M. Zhu et al., “A review of the application of machine learning in water quality evaluation,” Eco-Environment

Heal., vol. 1, no. 2, pp. 107–116, Jun. 2022, doi: 10.1016/j.eehl.2022.06.001.

U. Rohima Zalti, D. Rose Darmakusuma, M. Ridwansyah, and E. Ismanto, “Analisis dan Prediksi Kelayakan

Air Minum Menggunakan Algoritma Random Forest,” J. FASILKOM, vol. 15, no. 2, pp. 312–317, Aug. 2025,

doi: 10.37859/jf.v15i2.9906.

T. Z. Jasman, M. A. Fadhlullah, A. L. Pratama, and R. Rismayani, “Analisis Algoritma Gradient Boosting,

Adaboost dan Catboost dalam Klasifikasi Kualitas Air,” J. Tek. Inform. dan Sist. Inf., vol. 8, no. 2, Aug. 2022,

doi: 10.28932/jutisi.v8i2.4906.

M. Dava Maulana, A. Id Hadiana, and F. Rakhmat Umbara, “Algoritma Xgboost Untuk Klasifikasi Kualitas

Air Minum,” JATI (Jurnal Mhs. Tek. Inform., vol. 7, no. 5, pp. 3251–3256, 2024, doi: 10.36040/jati.v7i5.7308.

Prasetya Widiharso, Siti Sendari, Anik Nur Handayani, and Nastiti Susetyo Fanani Putri, “Performa Metode

Klasifikasi Tunggal dan Ensemble Model dalam Identifikasi Baku Mutu Air,” Infotekmesin, vol. 13, no. 2, pp.

–211, Jul. 2022, doi: 10.35970/infotekmesin.v13i2.1529.

B. M. Karomah, “PENERAPAN METODE STACKING DALAM MENGKLASIFIKASIKAN

PENDERITA PENYAKIT DIABETES,” J. Publ. Ilmu Komput. dan Multimed., vol. 1, no. 3, pp. 188–194, Jan. 1970, doi: 10.55606/jupikom.v1i3.522.

E. Prasetyo and K. Nugroho, “Optimasi Klasifikasi Data Stunting Melalui Ensemble Learning pada Label

Multiclass dengan Imbalance Data,” Techno.Com, vol. 23, no. 1, pp. 1–10, Feb. 2024, doi: 10.62411/tc.v23i1.9779.

M. Hermansyah, A. Saikhu, and B. Amaliah, “Pemodelan data radiosonde menggunakan stacking ensemble untuk klasifikasi hujan,” vol. 10, no. 2, pp. 1678–1687, 2025.

M. Kumar, S. Singhal, S. Shekhar, B. Sharma, and G. Srivastava, “Optimized Stacking Ensemble Learning

Model for Breast Cancer Detection and Classification Using Machine Learning,” Sustainability, vol. 14, no.

, p. 13998, Oct. 2022, doi: 10.3390/su142113998.

A. S. Alfath and A. K. Wardhana, “Hypertension Risk Prediction Using Stacking Ensemble of CatBoost ,

XGBoost , and LightGBM : A Machine Learning Approach,” vol. 9, no. 6, pp. 3146–3156, 2025.

F. M. Model, “Stacking Ensemble Learning : Combining XGBoost , LightGBM , CatBoost , and AdaBoost

with Random,” 2025.

A. A. Saputra, B. N. Sari, and C. Rozikin, “Penerapan Algoritma Extreme Gradient Boosting (Xgboost) Untuk

Analisis Risiko Kredit,” vol. 05, no. 01, pp. 53–64, 2022.

F. I. Kurniadi and P. D. Larasati, “Light Gradient Boosting Machine untuk Deteksi Penyakit Stroke,” J.

SISKOM-KB (Sistem Komput. dan Kecerdasan Buatan), vol. 6, no. 1, pp. 67–72, 2022, doi: 10.47970/siskomkb.v6i1.328.

A. Fahmi, L. Ptr, M. M. Siregar, and I. Daniel, “Journal of Computer Networks , Architecture and High

Performance Computing Analysis of Gradient Boosting , XGBoost , and CatBoost on Mobile Phone

Classification Journal of Computer Networks , Architecture and High Performance Computing,” vol. 6, no. 2, pp. 661–670, 2024.

D. Feby, “Serba Serbi Machine Learning Model Random Forest.” [Online]. Available: https://dqlab.id/serbaserbi-machine-learning-model-random-forest

A. D. Rachmatsyah, T. Sugihartono, and K. Irfan, “PERBANDINGAN TEKNIK OPTIMASI GRID

SEARCH DAN RANDOMIZED SEARCH DALAM MENINGKATKAN AKURASI METODE KLASIFIKASI SVM PADA SENTIMEN ULASAN PENGGUNA APLIKASI JKN MOBILE,” SKANIKA

Sist. Komput. dan Tek. Inform., vol. 8, no. 1, pp. 13–22, Dec. 2024, doi: 10.36080/skanika.v8i1.3328.

S. Q. Sultan, N. Javaid, and N. Alrajeh, “Machine Learning-Based Stacking Ensemble Model for Prediction

of Heart Disease with Explainable AI and K-Fold Cross-Validation : A Symmetric Approach,” no. Cvd, pp.

–26, 2025.

H. S. Silva et al., “The Impact of Feature Scaling In Machine Learning : Effects on Regression and

Classification Tasks,” 2025, doi: 10.1109/ACCESS.2025.3635541.

C. Paramita, C. S. Simbolon, A. S. Pamungkas, J. M. Triono, P. W. Utomo, and E. R. Subhiyakto, “Jurnal

Informatika : Jurnal pengembangan IT Analisis Pengaruh SMOTE terhadap Kinerja Model KNN untuk

Prediksi Risiko Stroke,” vol. 10, no. 4, pp. 978–988, 2025, doi: 10.30591/jpit.v10i4.8809.

K. S. Hasanah, Uswatun Agus Mohamad Soleh, “Effect of Random Under sampling , Oversampling , and

SMOTE on the Performance of Cardiovascular Disease Prediction Models terhadap Kinerja Model Prediksi

Penyakit Kardiovaskular,” vol. 21, no. 1, pp. 88–102, 2024, doi: 10.20956/j.v21i1.35552.

G. Erutjahjo and A. Supriyanto, “Jurnal Informatika : Jurnal pengembangan IT Prediksi Tinggi Gelombang

Laut di Perairan Semarang – Demak dengan Menggunakan Random Forest dan XGBoost,” vol. 10, no. 4, pp. 869–881, 2025, doi: 10.30591/jpit.v10i4.9315.




DOI: https://doi.org/10.30591/jpit.v11i2.10090

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.