Perbandingan Model Pohon Keputusan dan Random Forest dalam Mendeteksi Sentimen Isu Pajak Pertambahan Nilai

Janji Armanda Fabian, Galih Hendro Martono, Ismarmiaty Ismarmiaty

Abstract


Kenaikan PPN menjadi 12% memunculkan berbagai tanggapan dari masyarakat yang diekspresikan secara luas melalui media sosial, termasuk platform media sosial paling terkenal yaitu Youtube. Fenomena ini menunjukkan bahwa opini publik di dunia maya dapat menjadi bahan kajian penting dalam memahami respon masyarakat terhadap kebijakan fiskal pemerintah. Penelitian ini bertujuan untuk mengklasifikasikan emosi masyarakat terhadap isu kenaikan PPN 12% dengan memanfaatkan komentar YouTube sebagai sumber data. Tahapan penelitian ini dimulai dari pengumpulan data dengan metode web scraping, diikuti tahap pra-pemrosesan data, penerjemahan teks ke dalam bahasa Inggris, serta pelabelan emosi menggunakan NRC Lexicon. Selanjutnya, fitur teks diekstraksi menggunakann TF-IDF untuk kemudian dianalisis dengan Random Forest dan Decision Tree. Didapatkan hasil yang menunjukkan bahwa emosi dominan yang muncul pada komentar masyarakat adalah netral, positif, dan negatif. Model Decision Tree menghasilkan akurasi 57,35% dan macro F1-Score 0,53, sementara Random Forest memperoleh akurasi lebih tinggi sebesar 65,43% dengan macro F1-Score 0,57. Berdasarkan hasil, dapat dilihat Random Forest menunjukkan performa yang lebih tinggi daripada Decision Tree untuk klasifikasi teks.

Keywords


analisis sentimen; klasifikasi emosi; random forest; decision tree

Full Text:

References


E. Sihombing, M. Halmi Dar, and F. Aini Nasution, “Comparison of Machine Learning Algorithms in Public Sentiment Analysis of TAPERA Policy,” Int. J. Sci. Technol. Manag., vol. 5, no. 5, pp. 1089–1098, 2024, doi: 10.46729/ijstm.v5i5.1164.

N. S. I. Al-agele, “A Social Media Sentiment Analysis Using Machine Learning Approaches,” vol. 30, no. Ml, pp. 70–82, 2025.

K. Munger, J. Phillips, and T. George, “Pressing Play on Politics : Quantitative Description of YouTube James Bisbee Omer Yalcin University of Massachusetts Amherst , USA Cardiff University , Wales Matthew Hindman,” vol. 5, pp. 1–37, 2025.

P. Arsi and R. Waluyo, “Analisis Sentimen Wacana Pemindahan Ibu Kota Indonesia Menggunakan Algoritma Support Vector Machine (SVM),” J. Teknol. Inf. dan Ilmu Komput., vol. 8, no. 1, p. 147, 2021, doi: 10.25126/jtiik.0813944.

M. Taufiq Anwar, D. Riandhita Arief Permana, P. STMI Jakarta, P. Sistem Informasi Industri Otomotif, J. Letjen Suprapto No, and J. Pusat, “Analisis Sentimen Masyarakat Indonesia Terhadap Produk Kendaraan Listrik Menggunakan VADER,” Tek. Inform. dan Sist. Inf., vol. 10, no. 1, pp. 783–792, 2023, [Online]. Available: https://jurnal.mdp.ac.id/index.php/jatisi/article/view/3406/1173

W. Apriliah, I. Kurniawan, M. Baydhowi, and T. Haryati, “Prediksi Kemungkinan Diabetes pada Tahap Awal Menggunakan Algoritma Klasifikasi Random Forest,” Sistemasi, vol. 10, no. 1, p. 163, 2021, doi: 10.32520/stmsi.v10i1.1129.

A. S. Aribowo, H. Basiron, N. F. A. Yusof, and S. Khomsah, “Cross-domain sentiment analysis model on indonesian youtube comment,” Int. J. Adv. Intell. Informatics, vol. 7, no. 1, pp. 12–25, 2021, doi: 10.26555/ijain.v7i1.554.

C. Slamet, R. Andrian, D. S. Maylawati, Suhendar, W. Darmalaksana, and M. A. Ramdhani, “Web Scraping and Naïve Bayes Classification for Job Search Engine,” IOP Conf. Ser. Mater. Sci. Eng., vol. 288, no. 1, 2018, doi: 10.1088/1757-899X/288/1/012038.

Ф. Котлер et al., “No 主観的健康感を中心とした在宅高齢者における 健康関連指標に関する共分散構造分析Title,” Accid. Anal. Prev., vol. 183, no. 2, pp. 153–164, 2023.

Q. Alvi, S. F. Ali, S. B. Ahmed, N. A. Khan, M. Javed, and H. Nobanee, “On the frontiers of Twitter data and sentiment analysis in election prediction: a review,” PeerJ Comput. Sci., vol. 9, pp. 1–25, 2023, doi: 10.7717/peerj-cs.1517.

P. Mishra, A. Biancolillo, J. M. Roger, F. Marini, and D. N. Rutledge, “New data preprocessing trends based on ensemble of multiple preprocessing techniques,” TrAC - Trends Anal. Chem., vol. 132, p. 116045, 2020, doi: 10.1016/j.trac.2020.116045.

F. H. Nesa and A. C. Siregar, “Implementasi Data Mining untuk Klasifikasi Komentar Hate Speech Menggunakan Algoritma Support Vector Machine ( SVM ) Implementation of Data Mining for Hate Speech Comment Classification Using Support Vector Machine ( SVM ) Algorithm,” vol. 14, no. 105, 2025.

A. Chamekh, M. Mahfoudh, and G. Forestier, “Sentiment Analysis Based on Deep Learning in E-Commerce,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 13369 LNAI, pp. 498–507, 2022, doi: 10.1007/978-3-031-10986-7_40.

S. Astuti, A. Erfina, and U. N. Putra, “Analisis Sentimen Pandangan Publik Terhadap Kenaikan Pajak 12 % dari Twitter Menggunakan Indonesian Roberta Base Classifier Sentiment Analysis of Public Views on 12 % Tax Increase From Twitter Using English Roberta Base Classifier,” vol. 14, no. 105, pp. 698–706, 2025.

B. Khemani and A. Adgaonkar, “A Review on Reddit News Headlines with NLTK tool,” SSRN Electron. J., pp. 1–5, 2021, doi: 10.2139/ssrn.3834240.

A. A. Aliero, B. S. Adebayo, H. O. Aliyu, A. G. Tafida, B. U. Kangiwa, and N. M. Dankolo, “Systematic Review on Text Normalization Techniques and its Approach to Non-Standard Words,” Int. J. Comput. Appl., vol. 185, no. 33, pp. 44–55, 2023, doi: 10.5120/ijca2023923106.

S. Sarica and J. Luo, “Stopwords in technical language processing,” PLoS One, vol. 16, no. 8 August, pp. 1–13, 2021, doi: 10.1371/journal.pone.0254937.

S. Perveen et al., “Unsupervised fake news detection on social media using hybrid Gaussian Mixture Model,” PLoS One, vol. 20, no. 8 August, pp. 1–26, 2025, doi: 10.1371/journal.pone.0330421.

D. Ataman, A. Birch, N. Habash, M. Federico, and P. Koehn, “Machine Translation in the Era of Large Language Models : A Survey of Historical and Emerging Problems,” pp. 1–36, 2025.

A. Segura Navarrete, C. Martinez-Araneda, C. Vidal-Castro, and C. Rubio-Manzano, “A novel approach to the creation of a labelling lexicon for improving emotion analysis in text,” Electron. Libr., vol. 39, no. 1, pp. 118–136, 2021, doi: 10.1108/EL-04-2020-0110.

A. S. Ahmed, A. A. A. Haddad, R. S. Hameed, and M. S. Taha, “An Accurate Model for Text Document Classification Using Machine Learning Techniques,” Ing. des Syst. d’Information, vol. 30, no. 4, pp. 913–921, 2025, doi: 10.18280/isi.300408.

H. Luzern, Cyber Security Cyber Security, no. March. 2017. doi: 10.1007/978-981-16-9229-1.




DOI: https://doi.org/10.30591/smartcomp.v15i3.9776

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

========================================================================

Smart Comp Indexed By:

Flag Counter

 

View My Stats

 

 

 

Creative Commons License

 

 

 

This work is licensed under a Creative Commons Attribution 4.0 International License.