Evaluasi Gemini Flash pada Ekstraksi Jadwal Skripsi Terstruktur dan Tidak Terstruktur

Galih Hermawan, Ednawati Rainarli

Abstract


The administration of thesis seminar and defense scheduling is often hampered by unstructured PDF formats, which increases manual workload and the risk of errors. This study aims to evaluate and compare the performance of three Gemini Flash model variants, namely Gemini 2.0 Flash-Lite, Gemini 2.0 Flash, and Gemini 2.5 Flash Preview, in automating schedule information extraction using a zero-shot prompting approach. The dataset consists of 87 PDF files containing thesis seminar and defense schedules (588 entries) from the 2023/2024 academic year, alongside 200 question scenarios executed in two different context formats: raw extracted text (TXT) and structured JSON data. Performance evaluation metrics include Precision, Recall, F1-score, Exact-Match, and inference latency per request. Experimental results indicate that Gemini 2.5 Flash Preview achieves average F1-scores above 0.98 in both contexts with approximately 3.9 seconds latency. Conversely, smaller-capacity variants (Gemini 2.0 Flash and Flash-Lite) showed more significant performance gains using the JSON format compared to raw text, especially on complex question types such as multi-attribute filtering and list retrieval. Through error analysis, the primary challenge identified was tasks requiring numeric aggregation and determination of superlative values, accounting for approximately 78% of total extraction failures, particularly for lightweight models. A paired t-test indicated no statistically significant difference between the two context formats (average F1 difference = 0.0077; p=0.48). This study recommends the use of explicit numeric prompting or rule-based post-processing when employing lightweight models to significantly improve the accuracy of academic schedule information extraction.

Keywords


Ekstraksi Informasi; Gemini Flash; Jadwal Skripsi; Model Bahasa Besar; Zero-Shot Prompting

Full Text:

References


M. Arnold, M. Goldschmitt, and T. Rigotti, “Dealing with information overload: a comprehensive review,” Front. Psychol., vol. 14, June 2023, doi: 10.3389/fpsyg.2023.1122200.

S. Al Banna, I. Rakhmani, and D. C. Sukmajati, “Kilas Kebijakan PSPK – Membangun Keahlian dengan Profesionalisme di Kampus Merdeka: Kajian mengenai Beban Kerja Dosen di Indonesia,” Pusat Studi Pendidikan dan Kebijakan (PSPK), Jakarta, Apr. 2022. [Online]. Available: https://pspk.id/wp-content/uploads/2022/04/Kilas-Kebijakan-PSPK-Membangun-Keahlian-dengan-Profesionalisme-di-Kampus-Merdeka-Kajian-Mengenai-Beban-Kerja-Dosen-di-Indonesia.pdf

OECD, “Ensuring quality in VET and higher education: Getting quality assurance right,” OECD Education Policy Perspectives, Feb. 2025. doi: 10.1787/812ff006-en.

F. P. Diallo and C. Tudose, “Optimizing the Scheduling of Teaching Activities in a Faculty,” Appl. Sci., vol. 14, no. 20, Art. no. 20, Jan. 2024, doi: 10.3390/app14209554.

T. Brown et al., “Language Models are Few-Shot Learners,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf

D. Xu et al., “Large language models for generative information extraction: a survey,” Front. Comput. Sci., vol. 18, no. 6, p. 186357, Nov. 2024, doi: 10.1007/s11704-024-40555-y.

B. Grothey et al., “Comprehensive testing of large language models for extraction of structured data in pathology,” Commun. Med., vol. 5, no. 1, pp. 1–13, Mar. 2025, doi: 10.1038/s43856-025-00808-8.

V. Ntinopoulos et al., “Large language models for data extraction from unstructured and semi-structured electronic health records: a multiple model performance evaluation,” BMJ Health Care Inform., vol. 32, no. 1, p. e101139, Jan. 2025, doi: 10.1136/bmjhci-2024-101139.

A. Meddeb et al., “Evaluating local open-source large language models for data extraction from unstructured reports on mechanical thrombectomy in patients with ischemic stroke,” J. NeuroInterventional Surg., Aug. 2024, doi: 10.1136/jnis-2024-022078.

D. Jobson and Y. Li, “Investigating the Potential of Using Large Language Models for Scheduling,” in Proceedings of the 1st ACM International Conference on AI-Powered Software, in AIware 2024. New York, NY, USA: Association for Computing Machinery, July 2024, pp. 170–171. doi: 10.1145/3664646.3665084.

Y. Chai, Y. Fang, Q. Peng, and X. Li, “Tokenization Falling Short: On Subword Robustness in Large Language Models,” in Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds., Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 1582–1599. doi: 10.18653/v1/2024.findings-emnlp.86.

T. A. Chang and B. K. Bergen, “Language Model Behavior: A Comprehensive Survey,” Comput. Linguist., vol. 50, no. 1, pp. 293–350, Mar. 2024, doi: 10.1162/coli_a_00492.

D. Balsiger, H.-R. Dimmler, S. Egger-Horstmann, and T. Hanne, “Assessing Large Language Models Used for Extracting Table Information from Annual Financial Reports,” Computers, vol. 13, no. 10, Art. no. 10, Oct. 2024, doi: 10.3390/computers13100257.

Gemini Team and Google, “Gemini: A Family of Highly Capable Multimodal Models,” Google, 2024. Accessed: May 28, 2025. [Online]. Available: https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf

“Gemini 2.5: Our most intelligent models are getting even better,” Google. Accessed: June 09, 2025. [Online]. Available: https://blog.google/technology/google-deepmind/google-gemini-updates-io-2025/

V. N. Box Head of Product Marketing, Platform at, “Gemini 2.5 Flash Delivers Enhanced Document Q&A and Extraction with Box AI,” Box Blog. Accessed: June 09, 2025. [Online]. Available: https://backend.blog.box.com/gemini-25-flash-delivers-enhanced-document-qa-and-extraction-box-ai

P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, J. Su, K. Duh, and X. Carreras, Eds., Austin, Texas: Association for Computational Linguistics, Nov. 2016, pp. 2383–2392. doi: 10.18653/v1/D16-1264.

S. Krishna et al., “Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang, Eds., Albuquerque, New Mexico: Association for Computational Linguistics, Apr. 2025, pp. 4745–4759. Accessed: May 04, 2025. [Online]. Available: https://aclanthology.org/2025.naacl-long.243/

N. F. Liu et al., “Lost in the Middle: How Language Models Use Long Contexts,” Trans. Assoc. Comput. Linguist., vol. 12, pp. 157–173, 2024, doi: 10.1162/tacl_a_00638.

D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner, “DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio, Eds., Minneapolis, Minnesota: Association for Computational Linguistics, June 2019, pp. 2368–2378. doi: 10.18653/v1/N19-1246.

S. Neelam et al., “SYGMA: A System for Generalizable and Modular Question Answering Over Knowledge Bases,” in Findings of the Association for Computational Linguistics: EMNLP 2022, Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds., Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 3866–3879. doi: 10.18653/v1/2022.findings-emnlp.284.

G. Hermawan and E. Rainarli, “Evaluasi-LLM-Gemini-Jadwal-Skripsi — Dataset & Code.” GitHub repository, June 02, 2025. Accessed: June 10, 2025. [Online]. Available: https://github.com/Galih-Hermawan-Unikom/evaluasi-llm-gemini-jadwal-skripsi

L. N. Yaddanapudi, “The American Statistical Association statement on P-values explained,” J. Anaesthesiol. Clin. Pharmacol., vol. 32, no. 4, pp. 421–423, 2016, doi: 10.4103/0970-9185.194772.

D. Lakens, “Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs,” Front. Psychol., vol. 4, p. 863, Nov. 2013, doi: 10.3389/fpsyg.2013.00863.

P. S. P. Silveira, J. E. Vieira, and J. de O. Siqueira, “Is the Bland-Altman plot method useful without inferences for accuracy, precision, and agreement?,” Rev. Saúde Pública, vol. 58, p. 01, doi: 10.11606/s1518-8787.2024058005430.




DOI: https://doi.org/10.30591/jpit.v10i4.9047

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.

JPIT INDEXED BY

  
  

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.