Pipeline NLP End-to-End untuk Peringkasan Abstraktif dan Ekstraksi Entitas Berita Berbahasa Indonesia Berbasis Model Transformer
Abstract
The rapid growth of online news content poses challenges for readers to capture the core information quickly and accurately. This research proposes and implements an automated end-to-end pipeline that integrates three main stages: data acquisition, abstractive text summarization, and Named Entity Recognition (NER). The mT5 model is employed to generate coherent and concise summaries, while the BERT model is applied to extract key entities, including persons, organizations, and locations. The pipeline was evaluated using 100 news articles from the Egindo portal. Experimental results show that the system achieves an average text reduction of 62.47%, with a ROUGE-1 F1 score of 0.473. For NER tasks, the pipeline reached a Micro-F1 score close to 0.70, outperforming traditional approaches such as TextRank and CRF. These results demonstrate that the integration of Transformer-based models within a structured pipeline significantly improves summarization quality and entity extraction accuracy. The study contributes a practical NLP solution for the Indonesian language, providing a functional prototype that can be applied to online media analysis and media intelligence applications.
Keywords
References
A. Muharom and N. Rukhviyanti, “Development of Web-Based Multimedia Learning for Grade 3 Elementary School Mathematics,” INOVTEK Polbeng - Seri Informatika, vol. 10, no. 2, pp. 1142–1152, Jul. 2025, doi: 10.35314/sj1qng08.
K. Aggarwal, “A Review of Text Summarization Techniques Using NLP,” Computational Intelligence and Machine Learning, vol. 4, no. 2, Oct. 2023.
K. Bagla, A. Kumar, S. Gupta, and A. Gupta, “Noisy Text Data: Achilles’ Heel of Popular Transformer Based NLP Models,” Oct. 2021.
D. Nagalavi and M. Hanumanthappa, “The NLP Techniques for Automatic Multi-article News Summarization Based on Abstract Meaning Representation,” in Proceedings of the Conference, 2019, pp. 253–260.
S. Sistla, “Named Entity Recognition : A Deep Dive,” Journal of Artificial Intelligence & Cloud Computing, vol. 3, no. 6, pp. 1–5, Dec. 2024, doi: 10.47363/JAICC/2024(3)409.
A. V. Patil, “Identifying specific details from text to populate databases and generate summaries using Named Entity Recognition ,” INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT, vol. 08, no. 05, pp. 1–5, May 2024, doi: 10.55041/IJSREM33111.
S. Mhatre and L. Ragha, “Implementing Extractive Summarization Methods on Extractive Datasets,” in 5th International Conference on Data Intelligence and Cognitive Informatics (ICDICI), 2024, pp. 578–854.
G. Cheirmpos, S. A. Tabatabaei, E. Kanoulas, and G. Tsatsaronis, “Benchmarking Named Entity Recognition Approaches for Extracting Research Infrastructure Information from Text,” in Proceedings of the Conference, 2024, pp. 131–141.
F. M. Apriansyah, T. I. Ramadhan, C. R. Hidayat, and A. K. Wijaya, “Perbandingan IndoBERT dan IndoRoBERTa Untuk Analisis Sentimen Pada Film Dokumenter Dirty Vote,” Jurnal Informatika: Jurnal Pengembangan IT, vol. 10, no. 3, pp. 593–605, Jul. 2025, doi: 10.30591/jpit.v10i3.8607.
N. Widaningsih, N. Windiyanti, and N. Rukhviyanti, “Web-based Inventory Information System using Agile Scrum Method at CV Tunggal Putra Jaya,” SISTEMASI, vol. 14, no. 3, p. 1471, May 2025, doi: 10.32520/stmsi.v14i3.5253.
D. N. Pryatama and N. Rukhviyanti, “Rancang Bangun Aplikasi Stok Barang dengan QRcode Menggunakan Metode Waterfall dan Framwork Laravel pada Konveksi Sfgiandra,” JURNAL KRIDATAMA SAINS DAN TEKNOLOGI, vol. 7, no. 01, pp. 71–89, Feb. 2025, doi: 10.53863/kst.v7i01.1488.
K. V. Benedict and N. Rukhviyanti, “Analysis of the Classification of Data on the Launch of Apple Mobile Phone Prices in China and Pakistan Using the Decision Tree Algorithm in Python Programming,” Eduvest - Journal of Universal Studies, vol. 5, no. 9, pp. 10534–10546, Sep. 2025, doi: 10.59188/eduvest.v5i9.51409.
H. A. Holiel, N. Mohamed, A. Ahmed, and W. Medhat, “English-Arabic Text Translation and Abstractive Summarization Using Transformers,” in 20th ACS/IEEE International Conference on Computer Systems and Applications, 2023, pp. 1–8.
C. Meister, T. Vieira, and R. Cotterell, “Best-First Beam Search,” Trans. Assoc. Comput. Linguist., vol. 8, pp. 795–809, 2020, doi: 10.1162/tacl_a_00342.
N. Ott, R. Horst, and R. Dörner, “Towards Reducing Latency Using Beam Search in an Interactive Conversational Speech Agent,” in IEEE Gaming, Entertainment, and Media Conference (GEM), 2024, pp. 1–6.
S. Lemons, C. Linares López, R. C. Holte, and W. Ruml, “Beam Search: Faster and Monotonic,” in International Conference on Automated Planning and Scheduling (ICAPS), 2022, pp. 222–230.
M. W. A. Pramana, D. P. S. Putri, and I. K. A. Purnawan, “Comparison of IndoBERT and Bi-LSTM Models for Indonesian Law Violation Text Classification,” Jurnal Informatika: Jurnal Pengembangan IT, vol. 10, no. 4, pp. 1033–1043, Sep. 2025, doi: 10.30591/jpit.v10i4.8795.
Asro Asro, Meisa Monica, Novi Rukhviyanti, and M. Yusron, “Analisis Literatur Review Perencanaan Strategi Sistem Informasi Menggunakan Metode Pieces Framework,” KRESNA: Jurnal Riset dan Pengabdian Masyarakat, vol. 4, no. 2, pp. 161–169, Nov. 2024, doi: 10.36080/kresna.v4i2.182.
D. Zatnika and N. Rukhviyanti, “Penerapan Metode Forward Chaining pada Sistem Pakar Rekomendasi Mobil Second dari Aspek Penghasilan Kerja,” Jurnal Penelitian Inovatif, vol. 4, no. 4, pp. 2463–2476, Dec. 2024, doi: 10.54082/jupin.759.
T. Wolf and others, “Transformers: State-of-the-Art Natural Language Processing,” in EMNLP System Demonstrations, 2020, pp. 38–45. doi: 10.18653/v1/2020.emnlp-demos.6.
A. Barbaresi, “Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction,” in ACL-IJCNLP System Demonstrations, 2021, pp. 122–131. doi: 10.18653/v1/2021.acl-demo.15.
H. Tan, R. Yan, L. Yang, L. Huang, L. Xiao, and Q. Yang, “Efficient Multiple-Precision and Mixed-Precision Floating-Point Fused Multiply-Accumulate Unit for HPC and AI Applications,” in Proceedings of the Conference, 2023, pp. 642–659.
Susan Juli Safitri, Gelar Alam Ramdhaniawan, Asro Asro, and Novi Rukhviyanti, “Analisis Literatur Review Perencanaan Strategi Sistem Informasi Menggunakan Metode Metode Five Competitive Force Pada CV. Bio Chitosan Indonesia,” Bridge : Jurnal publikasi Sistem Informasi dan Telekomunikasi, vol. 2, no. 4, pp. 319–327, Sep. 2024, doi: 10.62951/bridge.v2i4.263.
T. Hasan and others, “XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages,” in Findings of ACL-IJCNLP, 2021, pp. 4693–4703. doi: 10.18653/v1/2021.findings-acl.413.
A. Al-Numai and A. Azmi, “LEMMA-ROUGE: An Evaluation Metric for Arabic Abstractive Text Summarization,” Indonesian Journal of Computer Science, vol. 12, no. 2, pp. 470–481, 2023, doi: 10.33022/ijcs.v12i2.330.
D. Arias, “Statistical Bias,” in Translational Sports Medicine, Elsevier, 2023, pp. 163–164.
N. Campolungo, T. Pasini, D. Emelin, and R. Navigli, “Reducing Disambiguation Biases in NMT by Leveraging Explicit Word Sense Information,” in NAACL-HLT, 2022, pp. 4824–4838.
T. Blevins and L. Zettlemoyer, “Moving Down the Long Tail of Word Sense Disambiguation with Gloss Informed Bi-encoders,” in ACL, 2020, pp. 1006–1017.
H. S. Yoon and others, “SMSMix: Sense-Maintained Sentence Mixup for Word Sense Disambiguation,” in EMNLP Findings, 2022, pp. 1493–1502.
H. Bast, M. Hertel, and N. Prange, “A Fair and In-Depth Evaluation of Existing End-to-End Entity Linking Systems,” in EMNLP, 2023, pp. 6659–6672. doi: 10.18653/v1/2023.emnlp-main.414.
K. Zaporojets, J. Deleu, Y. Jiang, T. Demeester, and C. Develder, “Towards Consistent Document-level Entity Linking,” in ACL Short Papers, 2022, pp. 778–784. doi: 10.18653/v1/2022.acl-short.98.
S. Dash, G. Rossiello, N. Mihindukulasooriya, S. Bagchi, and A. Gliozzo, “Open Knowledge Graphs Canonicalization using Variational Autoencoders,” in EMNLP, 2021, pp. 10379–10394. doi: 10.18653/v1/2021.emnlp-main.808.
A. Hamdi, A. Jean-Caurant, N. Sidère, M. Coustaty, and A. Doucet, “Assessing and Minimizing the Impact of OCR Quality on Named Entity Recognition ,” in Proceedings of the Conference, 2020, pp. 87–101.
Y. Cao and A. Yusup, “Chinese Electronic Medical Record Named Entity Recognition based on BERT-WWM-IDCNN-CRF,” in Dependable Systems and Their Applications (DSA), 2022, pp. 582–589.
J. C.-W. Lin, J. M.-T. Wu, Y. Shao, M. Pirouz, and B. Zhang, “A Latent Variable CRF Model for Labeling Prediction,” in Proceedings of the Conference, 2019, pp. 68–78.
DOI: https://doi.org/10.30591/jpit.v11i1.10030
Refbacks
- There are currently no refbacks.

This work is licensed under a Creative Commons Attribution 4.0 International License.
JPIT INDEXED BY
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | |

This work is licensed under a Creative Commons Attribution 4.0 International License.








