IMPLEMENTASI K-MEANS DAN ANALISIS SENTIMEN KRITIK SARAN BERBASIS NLP PADA DATA MONEV BBPSDMP KOMINFO MAKASSAR

##plugins.themes.academic_pro.article.main##

Syahril Akbar
Muhammad Faisal
Rizki Yusliana Bakti
Muhammad Syafaat
Andi Makbul Syamsuri
Muhyiddin AM Hayat
Andi Lukman Anas

Abstract

Manual analysis of large-scale and unstructured textual feedback data is often inefficient and subjective, thereby hindering data-driven decision-making. This study aims to design and implement an integrated analytical workflow to automatically filter, cluster, and classify feedback data consisting of criticisms and suggestions. The research employs a hybrid approach that begins with TF-IDF-based data filtering, followed by dimensionality reduction using Latent Semantic Analysis (LSA), and topic clustering through K-Means clustering optimized with the Silhouette Score. The resulting cluster labels are then used as training data to build a Multinomial Naive Bayes classification model. The results show that this workflow successfully identified two main thematic clusters, namely "Criticism and Expectations" and "Suggestions and Compliments", and the classification model achieved an overall accuracy of 91%. Although class imbalance affected the recall of the minority class (47%), the model demonstrated high precision (95%) for that class. It is concluded that this hybrid approach effectively transforms raw data into structured insights, and utilizing clustering results as training data is an efficient strategy for automating feedback categorization, providing a reliable tool for institutional analysis.

##plugins.themes.academic_pro.article.details##

How to Cite
Syahril Akbar, Muhammad Faisal, Rizki Yusliana Bakti, Muhammad Syafaat, Andi Makbul Syamsuri, Muhyiddin AM Hayat, & Andi Lukman Anas. (2025). IMPLEMENTASI K-MEANS DAN ANALISIS SENTIMEN KRITIK SARAN BERBASIS NLP PADA DATA MONEV BBPSDMP KOMINFO MAKASSAR. Jurnal Informatika Progres, 17(2), 36-43. Retrieved from https://jurnal.stmikprofesional.ac.id/index.php/Progress/article/view/432

References

[1] Jain, P. K., Pamula, R., & Srivastava, G. (2021). A systematic literature review on machine learning applications for consumer sentiment analysis using online reviews. Computer Science Review, 41, 100413. https://doi.org/10.1016/j.cosrev.2021.100413

[2] Egger, R., & Yu, J. (2022). A topic modeling comparison between LDA, NMF, Top2Vec, and BERTopic to demystify Twitter posts. Frontiers in Sociology, 7, 886498. https://doi.org/10.3389/fsoc.2022.886498

[3] Khomsah, S., & Aribowo, A. S. (2020). Model text-preprocessing komentar YouTube dalam bahasa Indonesia. Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), 4(4), 648–654. https://doi.org/10.29207/resti.v4i4.2035

[4] Hanan, D. A., Husodo, A. Y., & Rassy, R. P. (2025). Sentiment study of ChatGPT on Twitter data with hybrid K-Means and LSTM. MATRIK: Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer, 24(2). https://doi.org/10.30812/matrik.v24i2.4791

[5] Zikrina, S. A., & Fitriyani. (2025). Advancing hate speech detection in Indonesian language using graph neural networks and TF-IDF. Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), 9(1), 137–145. https://doi.org/10.29207/resti.v9i1.6179

[6] García-Jurado, A., Pérez-Barea, J. J., & Nova, R. (2021). A new approach to social entrepreneurship: A systematic review and meta-analysis. Sustainability, 13(5), 2754. https://doi.org/10.3390/su13052754

[7] Ahmed, M., Seraj, R., & Islam, S. M. S. (2020). The k-means algorithm: A comprehensive survey and performance evaluation. Electronics, 9(8), 1295. https://doi.org/10.3390/electronics9081295

[8] Fakhrezi, M. F., Rochim, A. F., & Nugraheni, D. M. K. (2023). Comparison of sentiment analysis methods based on accuracy value: Case study Twitter mentions of academic article. Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), 7(1), 161–167. https://doi.org/10.29207/resti.v7i1.4767

[9] Killamsetty, K., Sivasubramanian, D., Ramakrishnan, G., & Iyer, R. (2021). GLISTER: Generalization based data subset selection for efficient and robust learning. Proceedings of the AAAI Conference on Artificial Intelligence, 35(9), 8110–8118. https://doi.org/10.1609/aaai.v35i9.16988

[10] Priambodo, W., & Zuliarso, E. (2024). Kombinasi K-Means dan LSTM untuk deteksi black campaign di media sosial pada calon presiden Indonesia 2024. Jurnal Teknik Informatika (JUTIF), 5(2), 539–550. https://doi.org/10.52436/1.jutif.2024.5.2.1635

Most read articles by the same author(s)