Pemodelan Topik pada Komentar YouTube Arra: Komparasi LDA dan K-Means Menggunakan Fitur Leksikal dan Semantik

Siti Nuradilla, Sabrina Adnin Kamila, Latifah Zahra, Cici Suhaeni, Bagus Sartono

Abstract


YouTube has become a platform for sharing content, including positive material and stereotypes that often trigger debates. One noteworthy phenomenon is the video of Arra, a toddler known for her remarkable communication skills. This uniqueness has drawn significant attention and sparked debates about the mismatch between her age and cognitive development. The diverse comments on Arra’s videos reflect sharply differing perspectives among netizens, making manual analysis highly challenging. Therefore, it is important to examine the topics discussed by netizens to understand the dominant issues emerging in these discussions. Through this approach, the public can gain insights, and parents may receive valuable input regarding child-rearing practices. The main objective of this study is to explore the effectiveness of the two methods and their combinations of text representations in identifying key topics within comments by comparing the coherence performance of the models. This research applies topic modeling to analyze comments using two primary approaches: Latent Dirichlet Allocation (LDA) and K-Means clustering. The study involves data collection through comment crawling, followed by text preprocessing and text representation using TF-IDF and GloVe embeddings. LDA and K-Means are then used to identify dominant topics appearing in the comments. The results show that LDA with TF-IDF achieved the highest coherence score of 0.662, although the resulting topics were still difficult to interpret due to overlap. Meanwhile, K-Means with GloVe 100D yielded a slightly lower coherence score of 0.6538 but outperformed in terms of interpretability. Therefore, K-Means with GloVe 100D is considered a more balanced approach in terms of both coherence and topic readability.


Keywords


GloVe, K-Means, LDA, TF-IDF, Topic Modeling

Full Text:

References


kompasiana.com, “Mengenal Ara, Balita Viral di TikTok karena Hobi Deep Talk - Kompasiana.com.” Diakses: 10 April 2025. [Daring]. Tersedia pada: https://www.kompasiana.com/ahmadsidikx/669aa19ec925c4337655cc02/mengenal-ara-balita-viral-di-tiktok-karena-hobi-deep-talk

N. A. Rakhmawati, R. B. Waskitho, D. A. Rahman, dan M. F. A. U. Nuha, “Klasterisasi Topik Konten Channel YouTube Gaming Indonesia Menggunakan Latent Dirichlet Allocation,” JIEET, vol. 5, no. 2, hlm. 78–83, Des 2021, doi: 10.26740/jieet.v5n2.p78-83.

H. Jelodar dkk., “Latent Dirichlet Allocation (LDA) and Topic modeling: models, applications, a survey,” 6 Desember 2018, arXiv: arXiv:1711.04305. doi: 10.48550/arXiv.1711.04305.

G. Rosalinda, R. Santoso, dan P. Kartikasari, “PEMODELAN TOPIK ULASAN APLIKASI NETFLIX PADA GOOGLE PLAY STORE MENGGUNAKAN LATENT DIRICHLET ALLOCATION,” J.Gauss, vol. 11, no. 4, hlm. 554–561, Feb 2023, doi: 10.14710/j.gauss.11.4.554-561.

J. A. Lossio-Ventura, S. Gonzales, J. Morzan, H. Alatrista-Salas, T. Hernandez-Boussard, dan J. Bian, “Evaluation of clustering and topic modeling methods over health-related tweets and emails,” Artificial Intelligence in Medicine, vol. 117, hlm. 102096, Jul 2021, doi: 10.1016/j.artmed.2021.102096.

M. Das, S. Kamalanathan, dan P. J. A. Alphonse, “A Comparative Study on TF-IDF Feature Weighting Method and its Analysis using Unstructured Dataset,” International Conference on Computational Linguistics and Intelligent Systems, 2023, doi: 10.48550/arXiv.2308.04037.

N. Badri, F. Kboubi, dan A. H. Chaibi, “Combining FastText and GloVe Word Embedding for Offensive and Hate speech Text Detection,” Procedia Computer Science, vol. 207, hlm. 769–778, 2022, doi: 10.1016/j.procs.2022.09.132.

D. Andra dan A. B. Baizal, “E-commerce Recommender System Using PCA and K-Means Clustering,” J. RESTI (Rekayasa Sist. Teknol. Inf.), vol. 6, no. 1, hlm. 57–63, Feb 2022, doi: 10.29207/resti.v6i1.3782.

N. L. P. M. Putu, Ahmad Zuli Amrullah, dan Ismarmiaty, “Analisis Sentimen dan Pemodelan Topik Pariwisata Lombok Menggunakan Algoritma Naive Bayes dan Latent Dirichlet Allocation,” RESTI, vol. 5, no. 1, hlm. 123–131, Feb 2021, doi: 10.29207/resti.v5i1.2587.

R. Sabbagh dan F. Ameri, “A Framework Based on K-Means Clustering and Topic modeling for Analyzing Unstructured Manufacturing Capability Data,” Journal of Computing and Information Science in Engineering, vol. 20, no. 1, hlm. 011005, Feb 2020, doi: 10.1115/1.4044506.

S. Khomsah dan A. S. Aribowo, “Model Text-Preprocessing Komentar YouTube Dalam Bahasa Indonesia,” vol. 4, no. 4, 2020.

D. Rifaldi, Abdul Fadlil, dan Herman, “Teknik Preprocessing Pada Text Mining Menggunakan Data Tweet ‘Mental Health,’” Decode, vol. 3, no. 2, hlm. 161–171, Apr 2023, doi: 10.51454/decode.v3i2.131.

R. A. Naufal dan A. R. Pratama, “Analisis Sentimen terhadap Cyberbullying di Media Sosial dengan CrowdTangle,” AUTOMATA, 2023.

S. Akuma, T. Lubem, dan I. T. Adom, “Comparing Bag of Words and TF-IDF with different models for hate speech detection from live tweets,” Int. j. inf. tecnol., vol. 14, no. 7, hlm. 3629–3635, Des 2022, doi: 10.1007/s41870-022-01096-4.

Y. Fu dan Y. Yu, “Research on Text Representation Method Based on Improved TF-IDF,” J. Phys.: Conf. Ser., vol. 1486, no. 7, hlm. 072032, Apr 2020, doi: 10.1088/1742-6596/1486/7/072032.

I. Nyoman Prayana Trisna, N. Wayan Emmy Rosiana Dewi, dan M. Alam Pasirulloh, “Oversampling vs. undersampling in TF-IDF variations for imbalanced Indonesian short texts classification,” TELKOMNIKA, vol. 23, no. 2, hlm. 382, Apr 2025, doi: 10.12928/telkomnika.v23i2.26510.

M. A. Ikfini M dan E. B. Setiawan, “Topic Detection on Twitter using GloVe with Convolutional Neural Network and Gated Recurrent Unit,” bits, vol. 5, no. 2, Sep 2023, doi: 10.47065/bits.v5i2.4057.

T. F. Abdillah, H. Hasmawati, dan B. Bunyamin, “Comparison of TF-IDF and GloVe Word Embedding for Sentiment Analysis of 2024 Presidential Candidates,” bits, vol. 6, no. 2, hlm. 961–969, Sep 2024, doi: 10.47065/bits.v6i2.5668.

B. Bengfort, R. Bilbro, dan T. Ojeda, “Applied Text Analysis with Python,” O’Reilly Media, Inc., 2018.

D. M. Blei, A. Y. Ng, dan M. I. Jordan, “Latent Dirichlet Allocation,” Journal of Machine Learning Research, vol. 3, 2003.

D. Maier dkk., “Applying LDA Topic modeling in Communication Research: Toward a Valid and Reliable Methodology,” Communication Methods and Measures, vol. 12, no. 2–3, hlm. 93–118, Apr 2018, doi: 10.1080/19312458.2018.1430754.

M. Faisal, E. M. Zamzami, dan Sutarman, “Comparative Analysis of Inter-Centroid K-Means Performance using Euclidean Distance, Canberra Distance and Manhattan Distance,” J. Phys.: Conf. Ser., vol. 1566, no. 1, hlm. 012112, Jun 2020, doi: 10.1088/1742-6596/1566/1/012112.

R. Albalawi, T. H. Yeap, dan M. Benyoucef, “Using Topic modeling Methods for Short-Text Data: A Comparative Analysis,” Front. Artif. Intell., vol. 3, hlm. 42, Jul 2020, doi: 10.3389/frai.2020.00042.

D. Cline dan J. Ryan, “Exploring Coherence Metrics for Optimizing Topic Models of Humpback Song,” 2020.

Z. N. Muna, B. D. Setiawan, dan R. S. Perdana, “Penerapan Pemodelan Topik Komentar Melalui Media Sosial Twitter Menggunakan Latent Dirichlet Allocation (Studi Kasus: Pemerintah Kota Malang),” Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, 2024.




DOI: https://doi.org/10.30591/jpit.v10i3.8763

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.