Klasifikasi Pertanyaan Quora Menggunakan Metode Keyword-based dan Analisis Sentimen dengan ComplementNB

Alwan Adiuntoro, Aria Hendrawan

Abstract


Text classification is a fundamental task in Natural Language Processing (NLP) that supports the categorization of data based on predefined labels. This study aims to evaluate the effectiveness of keyword-based labeling and sentiment analysis methods for text classification using the Quora Questions dataset. The dataset comprises 16,921 samples with imbalanced class distribution, where the opinion category dominates, while the hypothetical category is a minority class. The labeling process utilized a keyword-based approach for the fact and hypothetical categories, while the opinion category was labeled using sentiment analysis with the Vader Lexicon library. TF-IDF was employed as the feature representation method, with two approaches explored: n-gram range tuning (1–3) and without tuning. ComplementNB, designed for handling imbalanced datasets, was utilized for classification, with a training-test split of 70:30. The results show that the approach without n-gram tuning achieved the highest accuracy of 93.89%, with zero variance in cross-validation. Evaluation revealed that ComplementNB effectively handles class imbalance, as demonstrated by high precision and recall in the minority class. This study demonstrates that a simple approach combining keyword-based labeling and sentiment analysis can be effectively implemented for category-based text classification tasks, particularly in platforms like Quora. These findings are relevant for similar applications requiring real-time text classification with minimal complexity.

Keywords


sentiment analysis; ComplementNB; text classification; keyword-based labeling; Quora

Full Text:

References


H. Ihsaniyati, S. Sarwoprasodjo, P. Muljono, and D. Gandasari, “The Use of Social Media for Development Communication and Social Change: A Review,” 2023. doi: 10.3390/su15032283.

Y. A. Ahmed, M. N. Ahmad, N. Ahmad, and N. H. Zakaria, “Social media for knowledge-sharing: A systematic literature review,” 2019. doi: 10.1016/j.tele.2018.01.015.

G. Wang, K. Gill, M. Mohanlal, H. Zheng, and B. Y. Zhao, “Wisdom in the social crowd: An analysis of Quora,” in WWW 2013 - Proceedings of the 22nd International Conference on World Wide Web, 2013.

S. D. Anggraeni, “Quora: Situs Komunitas Tanya Jawab Sebagai Medium Diskursus Ruang Publik,” Jurnal Sosia Logica, vol. 2, no. 1, pp. 1–14, 2023.

I. Dergaa, K. Chamari, P. Zmijewski, and H. Ben Saad, “From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing,” Biol Sport, vol. 40, no. 2, 2023, doi: 10.5114/BIOLSPORT.2023.125623.

Q. Li et al., “A Survey on Text Classification: From Traditional to Deep Learning,” 2022. doi: 10.1145/3495162.

M. Thangaraj and M. Sivakami, “Text classification techniques: A literature review,” Interdisciplinary Journal of Information, Knowledge, and Management, vol. 13, 2018, doi: 10.28945/4066.

D. Rogers, A. Preece, M. Innes, and I. Spasić, “Real-Time Text Classification of User-Generated Content on Social Media: Systematic Review,” 2022. doi: 10.1109/TCSS.2021.3120138.

J. Hartmann, J. Huppertz, C. Schamp, and M. Heitmann, “Comparing automated text classification methods,” International Journal of Research in Marketing, vol. 36, no. 1, 2019, doi: 10.1016/j.ijresmar.2018.09.009.

M. Chandra, A. Rodrigues, and J. George, “An Enhanced Deep Learning Model for Duplicate Question Detection on Quora Question pairs using Siamese LSTM,” in IEEE International Conference on Distributed Computing and Electrical Circuits and Electronics, ICDCECE 2022, 2022. doi: 10.1109/ICDCECE53908.2022.9792906.

N. Ahamed and S. Ahangama, “A Review of Classification of Insincere Questions in Quora Using Deep Learning Approaches,” in 2023 IEEE 17th International Conference on Industrial and Information Systems, ICIIS 2023 - Proceedings, 2023. doi: 10.1109/ICIIS58898.2023.10253595.

B. Marapelli, S. Kadiyala, and C. S. Potluri, “Performance Analysis and Classification of Class Imbalanced Dataset Using Complement Naive Bayes Approach,” in Proceedings of the ACCTHPA 2023 - Conference on Advanced Computing and Communication Technologies for High Performance Applications, 2023. doi: 10.1109/ACCTHPA57160.2023.10083369.

H. Florenci Tapikap, B. S. Djahi, and T. Widiastuti, “MAIL MENGGUNAKAN METODE TRANSFORMED COMPLEMENT NAÏVE BAYES (TCNB),” J-ICON, vol. 7, no. 1, 2019.

J. D. M. Rennie, L. Shih, J. Teevan, and D. Karger, “Tackling the Poor Assumptions of Naive Bayes Text Classifiers,” in Proceedings, Twentieth International Conference on Machine learning, 2003.

MR ADEPU RAJESH and DR TRYAMBAK HIWARKAR, “Exploring Preprocessing Techniques for Natural LanguageText: A Comprehensive Study Using Python Code,” international journal of engineering technology and management sciences, vol. 7, no. 5, 2023, doi: 10.46647/ijetms.2023.v07i05.047.

S. S. M. M. Rahman, K. B. M. B. Biplob, M. H. Rahman, K. Sarker, and T. Islam, “An investigation and evaluation of N-gram, TF-IDF and ensemble methods in sentiment classification,” in Lecture Notes of the Institute for Computer Sciences, Social-Informatics and Telecommunications Engineering, LNICST, 2020. doi: 10.1007/978-3-030-52856-0_31.

M. Bader-El-Den, E. Teitei, and T. Perry, “Biased Random Forest for Dealing with the Class Imbalance Problem,” IEEE Trans Neural Netw Learn Syst, vol. 30, no. 7, 2019, doi: 10.1109/TNNLS.2018.2878400.

K. R. M. Fernando and C. P. Tsokos, “Dynamically Weighted Balanced Loss: Class Imbalanced Learning and Confidence Calibration of Deep Neural Networks,” IEEE Trans Neural Netw Learn Syst, vol. 33, no. 7, 2022, doi: 10.1109/TNNLS.2020.3047335.

J. A. Prenner and R. Robbes, “Making the Most of Small Software Engineering Datasets with Modern Machine learning,” IEEE Transactions on Software Engineering, vol. 48, no. 12, 2022, doi: 10.1109/TSE.2021.3135465.

C. P. Chai, “Comparison of text preprocessing methods,” Nat Lang Eng, vol. 29, no. 3, 2023, doi: 10.1017/S1351324922000213.

A. W. Blocker and X. L. Meng, “The potential and perils of preprocessing: Building new foundations,” Bernoulli, vol. 19, no. 4, 2013, doi: 10.3150/13-BEJSP16.

E. A. Felix and S. P. Lee, “Systematic literature review of preprocessing techniques for imbalanced data,” 2019. doi: 10.1049/iet-sen.2018.5193.

Encyclopedia of Machine learning and Data Mining. 2017. doi: 10.1007/978-1-4899-7687-1.

S. Xu, “Bayesian Naïve Bayes classifiers to text classification,” J Inf Sci, vol. 44, no. 1, 2018, doi: 10.1177/0165551516677946.

H. A. Parhusip, B. Susanto, L. Linawati, S. Trihandaru, Y. Sardjono, and A. S. Mugirahayu, “Classification Breast Cancer Revisited with Machine learning,” International Journal of Data Science, vol. 1, no. 1, 2020, doi: 10.18517/ijods.1.1.42-50.2020.




DOI: https://doi.org/10.30591/jpit.v10i2.7965

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.