Perbandingan IndoBERT dan IndoRoBERTa Untuk Analisis Sentimen Pada Film Dokumenter Dirty Vote

Fadhel Muhammad Apriansyah, Teguh Ikhlas Ramadhan, Cepi Rahmat Hidayat, Anggito Karta Wijaya

Abstract


Sentiment analysis is a technique in Natural Language Processing (NLP) used to identify and categorize opinions or emotions in text. This study compares the performance of two Transformer-based models, IndoBERT and IndoRoBERTa, in analyzing sentiment toward the documentary film Dirty Vote. The research process includes data collection, text preprocessing, lexicon-based sentiment labeling, and model evaluation using K-Fold Cross-Validation. The results show that IndoBERT achieved an average accuracy of 99%, higher than IndoRoBERTa, which achieved 94%. IndoBERT also demonstrated better alignment with lexicon-based labeling in classifying positive, negative, and neutral sentiments. In terms of architecture, IndoBERT employs static masking, while IndoRoBERTa applies dynamic masking, leading to differences in the models' sensitivity to textual meaning. IndoBERT tends to provide more definitive classifications for opinions or strong criticisms, whereas IndoRoBERTa more frequently categorizes ambiguous comments as neutral sentiment. The conclusion of this study indicates that IndoBERT outperforms IndoRoBERTa in sentiment analysis of the documentary film Dirty Vote, both in terms of accuracy and consistency with lexicon-based labeling. These findings provide insights into the effectiveness of Transformer-based models for sentiment analysis in the Indonesian language and can serve as a reference for further NLP model development.

Keywords


Sentiment Analysis, IndoBERT, IndoRoBERTa, Dirty Vote, Electoral Fraud, Natural Language Processing

Full Text:

References


A. Aditia and A. P. Riyandi, “Pengertian Film Dokumenter: Definisi, Jenis dan Contohnya,” Kompas Entertainment, 17 Oktober 2022. [Online]. Available: https://entertainment.kompas.com/read/2022/10/17/173309466/pengertian-film-dokumenter-definisi-jenis-dan-contohnya

YouTube, “Dirty Vote - YouTube View Count,” 2024. [Online]. Available: https://www.youtube.com

Google Trends, “Tren Pencarian ‘Dirty Vote’ di Indonesia (Februari 2024),” 2024. [Online]. Available: https://trends.google.com

K. Munger and J. Phillips, “A Supply and Demand Framework for YouTube Politics,” Public Opin. Q., vol. 86, no. S1, pp. 161–188, 2022. [Online]. Available: https://doi.org/10.1093/poq/nfac002

B. Pang and L. Lee, “Opinion Mining and Sentiment Analysis,” Found. Trends Inf. Retr., vol. 2, no. 1–2, pp. 1–135, 2008. [Online]. Available: https://doi.org/10.1561/1500000001

E. Cambria, B. Schuller, Y. Xia, and C. Havasi, “New Avenues in Opinion Mining and Sentiment Analysis,” IEEE Intell. Syst., vol. 28, no. 2, pp. 15–21, 2013. [Online]. Available: https://doi.org/10.1109/MIS.2013.30

B. Liu, Sentiment Analysis and Opinion Mining, vol. 5, no. 1. Synth. Lect. Hum. Lang. Technol., 2012, pp. 1–167. [Online]. Available: https://doi.org/10.2200/S00416ED1V01Y201204HLT016

R. Feldman, “Techniques and Applications for Sentiment Analysis,” Commun. ACM, vol. 56, no. 4, pp. 82–89, 2013. [Online]. Available: https://doi.org/10.1145/2436256.2436274

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint arXiv:1810.04805, 2019. [Online]. Available: https://doi.org/10.48550/arXiv.1810.04805

B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” arXiv preprint arXiv:2009.05387, 2020. [Online]. Available: https://doi.org/10.48550/arXiv.2009.05387

J. U. S. Lazuardi and A. Juarna, “Analisis Sentimen Ulasan Pengguna Aplikasi JOOX pada Android Menggunakan Metode BERT,” J. Ilm. Inform. Komput., vol. 28, no. 3, pp. 251–260, 2023. [Online]. Available: https://doi.org/10.35760/ik.2023.v28i3.10090

N. A. R. Putri and Ardiansyah, “Analisis Sentimen terhadap Kemajuan Kecerdasan Buatan di Indonesia Menggunakan BERT dan RoBERTa,” J. Sains Inform., vol. 9, no. 2, pp. 136–145, 2023. [Online]. Available: https://doi.org/10.34128/jsi.v9i2.649

E. P. A. Akhmad, “Analisis Sentimen Ulasan Aplikasi DLU Ferry pada Google Play Store Menggunakan BERT,” J. Apl. Pelayaran Kepelabuhanan, vol. 13, no. 2, pp. 104–112, 2023. [Online]. Available: https://doi.org/10.30649/japk.v13i2.94

K. Nahar, A. Jaradat, M. Atoum, and F. Ibrahim, “Sentiment Analysis and Classification of Arab Jordanian Facebook Comments Using Machine Learning,” Jordanian J. Comput. Inf. Technol., vol. 0, no. 1, p. 1, 2020. [Online]. Available: https://doi.org/10.5455/jjcit.71-1586289399

K. Koto et al., “IndoBERT: A Pre-trained Indonesian-Specific Language Model,” in Proc. 28th Int. Conf. Comput. Linguist. (COLING), 2020.

F. Pratama and H. Wibowo, “Improving Indonesian Sentiment Analysis Using IndoRoBERTa,” Indones. J. Comput. Cybern. Syst., vol. 15, no. 3, pp. 198–210, 2021.

A. S. Ramadhan, S. Wibowo, and M. Kurniawan, “Comparing IndoBERT, LSTM, and SVM for Sentiment Analysis on Indonesian News Data,” J. Inf. Syst. Eng. Bus. Intell., vol. 9, no. 1, pp. 23–31, 2022.

M. Runimeirati, A. Muis, and F. Muhammad, “Pelatihan Text Mining Menggunakan Bahasa Pemrograman Python,” Abdimas Langkanae, vol. 3, no. 1, pp. 36–46, 2023. [Online]. Available: https://doi.org/10.53769/abdimas.3.1.2023.83

A. Wijaya, “Comparative Analysis of Transformer-based Models for Indonesian Sentiment Analysis,” Int. J. Artif. Intell. Res., vol. 10, no. 2, pp. 87–95, 2023.




DOI: https://doi.org/10.30591/jpit.v10i3.8607

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.